LoopGain: stop your agent loop the moment it actually converges — hands-on with the control-theoretic cost controller
Every production agent loop ends the same way: max_iterations = N. LoopGain — the Apache-2.0 Python library climbing GitHub's trending lists this week — replaces that fixed cap with a control-theoretic stop-and-rollback policy: it watches your loop's error trajectory, stops the moment it actually converges, and hands back the best iteration when the loop degrades. Pure Python, zero dependencies, $0 to learn. Every command and code block below ran exactly as shown.
Why this is blowing up now
Every production agent loop you have written ends the same way: max_iterations = N. It is the embarrassing default of agentic AI — the project's README says so outright. Set the cap too low and you ship bad output; set it too high and you burn tokens on revisions that change nothing. Patience counters ("stop after three non-improving rounds") are only slightly better: they count iterations, not improvement.
LoopGain (loopgain-ai/loopgain, Apache-2.0) replaces the fixed cap with a control-theoretic stop-and-rollback policy grounded in the Barkhausen criterion — a 1921 result from feedback-oscillator analysis. It watches a single number, your loop's error signal, classifies the error trajectory with four statistical features, and stops the loop the moment it has actually converged. When the loop degrades instead of improving, it aborts and hands you back the best iteration it saw, not the wreckage.
It is surfacing on GitHub trending roundups this week — it led the September 23 "7 Open-Source AI Repos Worth Watching" list at 124 stars and climbing — and the numbers behind it are unusually well-documented for a young project. Its public benchmark (2,000 paired trials across 10 workload cells, reproducible from a companion repo) claims 92.8% less API spend than max_iter=20 ($27.05 down to $1.94 in total benchmark spend), roughly 15× faster median wall-clock (30.9s to 2.1s), judge win-rates of 0.50–0.63 on natural workloads and 0.92–0.95 on engineered-failure workloads — with zero of six pre-registered kill criteria firing. Treat those as the project's own benchmark claims, not independent verification.
What makes it worth a tutorial: it is pure Python with zero runtime dependencies, installs in seconds, costs $0 and needs no API key to learn, and its core API is three lines. Everything below ran exactly as shown, on loopgain 0.6.5.
What you'll need
- Python 3.10 or newer — check with
python3 --version. LoopGain supports 3.10+. pip, ideally inside a virtual environment. The library has no runtime dependencies.- A terminal. Nothing else: no API key, no GPU, no model weights, no account. Every example here runs with $0 of API spend, because the error signal is yours to define — I demonstrate it with deterministic trajectories you can inspect line by line.
- Optional, for the adapter section: one of LangGraph, CrewAI, AutoGen, LangChain, the OpenAI Agents SDK, or the Claude Agent SDK in your own project.
Step 1 — Install and verify
pip install loopgain
loopgain --version
Expected output (version at time of writing):
loopgain 0.6.5
Confirm the import works, and check the telemetry default — the library's anonymous usage telemetry is opt-in and default-decline:
python3 -c "from loopgain import LoopGain; print('import ok')"
loopgain telemetry --show
telemetry --show prints the current status and exactly what would be sent if you ever opted in. Nothing leaves your machine unless you explicitly enable it (loopgain telemetry --enable or LOOPGAIN_TELEMETRY=1); DO_NOT_TRACK=1 is honored as a hard opt-out, and CI environments are auto-declined silently. The fleet-dashboard telemetry is a separate, per-call opt-in covered in Step 6. The library itself never phones home.
Step 2 — Wrap your first loop
The whole core API is three calls: construct a monitor, observe() each iteration's error, and let should_continue() drive the loop:
from loopgain import LoopGain
lg = LoopGain(target_error=0.0, max_iterations=20)
errors = [12, 9, 7, 5, 3, 2, 1, 0] # your verifier's error per iteration
i = 0
while lg.should_continue() and i < len(errors):
state = lg.observe(errors[i], output=f"draft-v{i}")
print(f"iter {i}: error={errors[i]} -> {state}")
i += 1
r = lg.result
print(r.outcome, r.iterations_used, r.best_index, r.savings_vs_fixed_cap)
Run it. The exact output on 0.6.5:
iter 0: error=12 -> FAST_CONVERGE
iter 1: error=9 -> CONVERGING
iter 2: error=7 -> CONVERGING
iter 3: error=5 -> CONVERGING
iter 4: error=3 -> CONVERGING
iter 5: error=2 -> CONVERGING
iter 6: error=1 -> FAST_CONVERGE
iter 7: error=0 -> TARGET_MET
converged 8 7 2
Two things to notice. First, the short-circuit: with the default target_error=0.0, the moment an observation reads exactly zero, the loop stops with state TARGET_MET — zero is the natural completion signal for verifier-driven loops. Second, the result: outcome is "converged", iterations_used is 8, best_index is 7 (the zero-error iteration), and savings_vs_fixed_cap is 2 — the monitor assumed a fixed cap of 10 (the assumed_fixed_cap default) and you used 8. Pass assumed_fixed_cap=20 and the identical run reports 16.

Step 3 — Define your error signal
LoopGain knows nothing about your loop's domain. The one thing you provide is a single non-negative number per iteration that says how wrong the current output is — lower is better, zero means done. The project's docs suggest this mapping, worth stealing verbatim:
| Your loop | Error signal |
|---|---|
| Agentic coding (write code → run tests) | number of failing tests (10 → 3 → 0) |
| JSON / structured extraction | number of schema violations |
| RAG with self-correction | number of required facts still missing |
| Self-refinement with an LLM judge | judge's gap to target (e.g. 10 − quality_score) |
| Lint / format loop | lint error count |
Two mechanics matter.
Numbers or sequences. observe() accepts a plain number, or any sequence — in which case its length becomes the magnitude. Hand it the raw list of failing tests and never pre-count:
lg = LoopGain(target_error=0.0)
for failing, out in [(["t_login", "t_checkout", "t_refund"], "v0"),
(["t_refund"], "v1"),
([], "v2")]:
print(lg.observe(failing, output=out), lg.should_continue())
FAST_CONVERGE True
CONVERGING True
TARGET_MET False
An empty list has length zero, which trips the default target's short-circuit. The loop ends the moment the list is empty — which is exactly when you'd stop anyway.
When quality has no natural zero. Fuzzy targets — style scores, open-ended rewrites — never read exactly zero. Pass target_error=None and the short-circuit is disabled: the monitor then stops when the number stops improving, wherever the plateau sits. The next section shows exactly what that looks like.
Step 4 — End to end: a self-healing JSON repair loop
Now a complete loop with a real verifier. The scenario: an extractor produces JSON, a schema check counts violations, a reviser repairs them. To keep this runnable for $0, the reviser below is a deterministic stand-in for an LLM call — swap in your real model call and nothing else changes, because the only contract LoopGain sees is the error number:
from loopgain import LoopGain
REQUIRED = {"name": str, "email": str, "age": int}
def verify(payload):
"""Real verifier: returns the list of schema violations."""
problems = []
for key, typ in REQUIRED.items():
if key not in payload:
problems.append(f"missing:{key}")
elif not isinstance(payload[key], typ):
problems.append(f"type:{key}")
return problems
# Stand-in reviser: fixes one violation per iteration (your LLM goes here).
FIXES = [("name", "Ada"), ("email", "[email protected]"), ("age", 36)]
lg = LoopGain(target_error=0.0, max_iterations=20)
payload = {}
for i, (key, val) in enumerate(FIXES):
problems = verify(payload)
state = lg.observe(problems, output=dict(payload))
print(f"iter {i}: {len(problems)} violations -> {state}")
if not lg.should_continue():
break
payload[key] = val
else:
# fixes exhausted with the loop still alive: check the final payload
lg.observe(verify(payload), output=dict(payload))
r = lg.result
print("outcome:", r.outcome)
print("best output:", r.best_output)
print("iterations used:", r.iterations_used)
iter 0: 3 violations -> FAST_CONVERGE
iter 1: 2 violations -> CONVERGING
iter 2: 1 violations -> CONVERGING
outcome: converged
best output: {'name': 'Ada', 'email': '[email protected]', 'age': 36}
iterations used: 4
The verifier is doing the real work here — the list of violations is a genuine error signal, and the final observe() sees zero violations and fires TARGET_MET. best_output is the fully repaired payload.
Now the two cases that justify the library. First, the plateau: the loop improves, then stalls at 4 violations with target_error=None (no natural zero):
lg = LoopGain(target_error=None, max_iterations=20)
traj = [12, 10, 8, 6, 5, 4, 4, 4, 4, 4]
i = 0
while lg.should_continue() and i < len(traj):
print(f"iter {i}: error={traj[i]} -> {lg.observe(traj[i], output=f'v{i}')}")
i += 1
r = lg.result
print(r.outcome, r.iterations_used, r.best_index, r.best_error)
iter 0: error=12 -> FAST_CONVERGE
iter 1: error=10 -> CONVERGING
iter 2: error=8 -> CONVERGING
iter 3: error=6 -> CONVERGING
iter 4: error=5 -> CONVERGING
iter 5: error=4 -> CONVERGING
iter 6: error=4 -> CONVERGING
iter 7: error=4 -> CONVERGING
iter 8: error=4 -> STALLING
iter 9: error=4 -> STALLING
stalled 10 5 4.0
After two consecutive stalling readings the monitor stops and reports stalled — best error 4.0 at iteration 5, the first time the loop reached it. A fixed cap of 20 would have burned 10 more iterations for zero improvement. Note the discipline this forces on you: a stall at 4 violations is a quality gap the controller cannot see — check result.best_error before you trust the output.
Second, the divergence — the case that earns the "rollback" in the name. Error jumps upward instead of improving:
lg = LoopGain(target_error=None, max_iterations=25)
for i, e in enumerate([3, 4, 6, 9, 13, 18]):
state = lg.observe(e, output=f"v{i}")
print(f"iter {i}: error={e} -> {state}")
if not lg.should_continue():
break
r = lg.result
print(r.outcome, "| best:", r.best_output, "| iters:", r.iterations_used)
iter 0: error=3 -> FAST_CONVERGE
iter 1: error=4 -> DIVERGING
diverged | best: v0 | iters: 2
The monitor fired DIVERGING after a single upward jump, aborted the loop at iteration 2, and rolled back to v0 — the best output it had seen. Divergence detection becomes "abort with the best you've seen so far" instead of "abort with garbage." That is the free quality floor.

Step 5 — Read the result like a dashboard
After the loop, lg.result is a LoopGainResult. The fields you will actually use:
outcome— one ofconverged(target met or clean convergence),stalled(improvement plateaued),oscillating(bouncing with no trend),diverged(getting worse — aborted), ormax_iterations(the hard backstop fired).best_output/best_index/best_error— the lowest-error iteration observed. Ship this, not the last iteration.iterations_used,error_history(every observed error),convergence_profile, andsavings_vs_fixed_cap— which is simplyassumed_fixed_cap − iterations_used. Setassumed_fixed_capto your current fixed cap and the number becomes a direct apples-to-apples comparison with the policy you have today.
Behind the outcomes sits a five-state classifier. Each trajectory is routed on four features: cumulative error reduction, the OLS slope of log10(error), the slope's t-test p-value, and the residual oscillation magnitude. The decision is conservative by design — it requires both statistical significance and meaningful motion before terminating, which is why thin evidence (4 or fewer observations) falls back to stalling rather than risking a false abort. The project recommends a minimum of 6 iterations for reliable trend significance.
Two honest limits, stated in the project's own docs and worth repeating. LoopGain detects convergence, not correctness — it knows when more iterations won't help, not whether the answer is right. And it is only as good as your verifier: in the project's benchmark, 4.5% of converged runs (16 of 355) passed every in-loop check yet failed the full held-out test suite. Pair it with the strongest verifier you can afford at the stop — executable tests over a sampled subset, a schema check over a vibe, a held-out check the loop didn't optimize against.
Step 6 — Wire it into a real framework
If your loop already lives in an agent framework, you don't rewrite it. Thin adapters under loopgain.integrations drive the framework's own iteration with your monitor — the frameworks are optional dependencies, installed as extras:
pip install 'loopgain[langgraph]' # or crewai, autogen, langchain,
pip install 'loopgain[openai-agents]' # claude-agent-sdk, or [all]
Every adapter takes your LoopGain instance plus an error_fn you provide — the framework can't know what "wrong" means in your domain, so the adapter doesn't guess. The LangGraph shape, from the project's docs:
from loopgain import LoopGain
from loopgain.integrations import LangGraphAdapter
lg = LoopGain(target_error=0.1, max_iterations=20)
adapter = LangGraphAdapter(
lg=lg,
error_fn=lambda update: len(update.get("verifier", {}).get("errors", [])),
)
final_state = adapter.run(graph, {"draft": initial})
Adapters exist for CrewAI (task_error_fn / step_error_fn), AutoGen v0.4 (observe_sources={"verifier"}), LangChain, the OpenAI Agents SDK, and the Claude Agent SDK — all with the same lg + error_fn shape. I did not execute the adapter examples here (they need a live framework and model calls); the raw API in Steps 2–5 is the part verified end to end.
One more opt-in worth knowing: lg.send_telemetry(...) posts a single anonymized aggregate to your fleet dashboard after the loop terminates — state transitions, iteration counts, savings, never prompts, outputs, or error contents — and only when you call it. The hosted receiver and dashboard are both open-source and self-hostable. This is separate from the anonymous usage funnel in Step 1, which stays off unless you turn it on.
When to use this vs the alternatives
vs max_iterations: the fixed cap is either wasteful or dangerous, and you don't get to choose which on any given run. LoopGain makes the decision from the full error trajectory instead of the iteration count. The project's benchmark claims 92.8% spend savings against a cap of 20; your number depends on how often your loops currently overshoot.
vs a hand-rolled patience counter ("stop after k non-improving rounds"): patience counts flat iterations; LoopGain measures the trend statistically — it keeps going through noisy-but-improving runs a patience counter would kill, catches divergence a counter would sleep through, and hands you the rollback for free.
vs LLM-as-a-judge gating: a judge answers "is this good enough" per output; LoopGain answers "is more iterating worth it" across outputs. They compose: feed the judge's gap-to-target in as your error signal and LoopGain decides when to stop paying the judge.
vs token observability tools: dashboards like Agent Console tell you where the money went; LoopGain stops the spending. They compose too — measure with the dashboard, cap with the controller. The same bill is attacked from the prompt side by prompt compression, cascade routing, and RTK's output filtering; LoopGain is the only one of the four that watches the loop itself.
When not to use it: if your loop has no measurable error signal — pure open-ended generation with no verifier — there is nothing for the classifier to watch, and target_error=None on noise is just a slower way to hit max_iterations. Build the verifier first. And per the docs, don't tune TrajectoryThresholds until you have production traces of your own.
The takeaway
LoopGain does one thing: it watches the error trajectory of an iterative agent loop and stops it the moment more iterations stop helping — rolling back to the best version when the loop degrades. Three lines wrap any loop, the error signal is whatever your verifier already measures, and the whole thing is pure Python with no dependencies and no telemetry unless you opt in. Point it at your most expensive verify-revise loop with assumed_fixed_cap set to your current cap, and let savings_vs_fixed_cap tell you whether your fixed cap was ever the right policy. For most loops, it wasn't.