Loop engineering is the practice of designing the repeated cycle an agent runs through — think, act, observe, repeat — so that it makes real progress, stops at the right time, and fails safely when it does not. The loop is the heartbeat of most agentic systems, though some use fixed workflows or graphs instead. Small design mistakes in a loop can cause big problems.
At its simplest, an agent loop looks like this:
while task is not done:
decide next action
take the action
observe the result
That single while line hides most of the hard engineering work. Loop engineering is about answering: how does the agent decide it is “done”? What happens if it never is? What happens if it repeats the same failed step forever?
The pseudocode is simplified. Production loops also stop or pause on cancellation, tool errors, human approval waits, deadlines, and budget limits.
Why loops are risky by default
A loop with no limits will run until something forces it to stop. In an agentic system, that “something” needs to be designed deliberately, or a few things can go wrong:
- Infinite loops: the model keeps trying the same failing action, or oscillates between two actions without progress.
- Runaway cost: every iteration calls a model and possibly external tools, so an unbounded loop can become expensive quickly.
- Silent drift: the agent keeps “making progress” in a direction that is no longer useful, without anyone noticing.
- Compounding errors: a small mistake early in the loop feeds into every later step, growing worse each time.
Loop engineering exists to catch these problems before they cause damage.
Termination conditions
A well-engineered loop needs more than one way to stop:
- Externally verified success: the goal is met — for example, all tests pass, a record exists, or the user confirms the answer is correct.
- Model-reported completion: the model says it is done. This is useful, but it should be checked against external evidence when possible.
- Step limit: a maximum number of iterations, so the loop cannot run forever even if it never technically succeeds.
- Time or cost budget: a ceiling on how long or how much the loop is allowed to spend.
- No-progress detection: if the last few steps produced no meaningful change, the loop should stop and report rather than keep repeating.
- Stuck or failure signal: the model can report that it is stuck, and the harness should also detect stuck behavior independently.
Relying on just one of these (usually “the model says it’s done”) is fragile. Combining several makes the loop far more dependable, and those checks are usually enforced by the surrounding harness.
Detecting progress, not just activity
A loop can look busy while going nowhere — calling tools, producing output, taking actions — without moving closer to the goal. Good loop engineering distinguishes activity from progress.
One approach is to track a measurable signal tied to the actual goal. For a coding agent fixing a bug, that might be “number of failing tests.” For a research agent, it might be “number of open questions answered.” If the signal is not improving after a few iterations, the loop should reconsider its approach instead of repeating it.
Progress should mean a measurable improvement: failing tests drop, a new fact is found, a blocker is removed, a file changes in the intended area, or a user-approved checkpoint is reached. False progress is activity without improvement: repeated identical tool calls, rewriting the same paragraph, searching the same logs, or producing longer plans without new evidence.
Feeding the right context back in
Each turn of the loop, the harness has to decide what the model sees next — the original goal, what has happened so far, and the latest result. Including too little makes the model repeat past mistakes. Including too much (like the full history of every step) can overwhelm the context window and slow the model down. This connects closely to context engineering: loop design and context design are solved together, not separately.
Parallel steps add another risk: stale state. If two branches read the same state and both act on it, one branch may make the other’s plan invalid. A loop that runs parallel work should merge observations carefully, re-check assumptions, and avoid write actions based on stale data.
Failure handling inside the loop
Real loops need more than “try again.” Common controls include:
- Cancellation: stop promptly when the user cancels or the parent workflow ends.
- Human waits: pause with saved state when approval or extra input is needed, then resume from the same checkpoint.
- Tool timeouts: stop waiting on a slow tool and return a clear error to the loop.
- Retries with backoff: retry transient failures after a delay, but only when the action is safe to repeat.
- Circuit breakers: stop calling a failing tool after N failures and choose another path or report the blocker.
These controls keep a temporary service failure from becoming an expensive runaway loop.
An example: a research loop
Goal: Answer "What caused the outage last Tuesday?"
Iteration 1: Search incident logs. -> Found a timestamp, no root cause yet.
Iteration 2: Search deploys around that timestamp. -> Found a matching deploy.
Iteration 3: Read the deploy's change log. -> Found the likely change.
Iteration 4: Confirm by checking error rates before/after the change.
Iteration 5: Summarize findings and stop - goal met.
Each iteration should narrow the search or add new evidence. If iteration 6, 7, and 8 kept searching the same logs without new findings, a well-engineered loop would recognize the stall and stop rather than continue indefinitely, or hand the task to a human reviewer.
Loop metrics
Useful loop metrics include:
- Success rate: how often the loop reaches an externally verified goal.
- Premature-stop rate: how often it stops before enough work was done.
- Runaway-loop rate: how often it hits a hard limit instead of a natural stop.
- Average steps and cost per task: how much work successful runs take.
- Tool failure rate: how often retries, timeouts, or circuit breakers trigger.
These metrics make loop changes testable. A new prompt, tool, or stop rule should improve at least one metric without making another important one worse.
The key idea
Loop engineering is about making the repeated think-act-observe cycle of an agent reliable: knowing when to stop, recognizing real progress versus busywork, and failing safely instead of running forever. A good loop is what keeps an agentic system from spinning out of control.