A harness is the surrounding software that runs an AI agent: it manages the agent loop, calls tools, tracks state, enforces limits, and decides what the model sees at each step through context engineering. Harness engineering is the work of designing that surrounding system, as distinct from prompt engineering, which focuses on what you say to the model.
You can think of the model as the engine and the harness as everything else in the car — the frame, the controls, the brakes, and the dashboard. A powerful engine with no harness is just a risky, undirected source of power.
Why the model alone is not enough
A language model maps its input to output during inference. That output may be text, a structured tool call, a refusal, or an error-like partial response. The model does not hold durable state, run tools, enforce permissions, or know when to stop by itself. Memory features in products are implemented around the model.
The harness is responsible for things like:
- Storing and updating conversation history and task state.
- Calling tools when the model requests them and returning results.
- Deciding what context to include in the next model call.
- Detecting when the task is complete, stuck, or has failed.
- Enforcing limits on time, cost, and number of steps.
- Logging what happened for debugging and review.
None of this is the model’s job. It is the harness’s job.
The core loop
Most agent harnesses run some version of this loop:
1. Build the next input (instructions + relevant context + history).
2. Call the model.
3. Parse the response: final text, one or more tool calls, refusal, error, or partial output.
4. If a tool call: run it, capture the result, go back to step 1.
5. If a final answer: return it to the user (or end the task).
The loop itself sounds simple, but the quality of a harness usually comes down to the details around it, not the loop structure.
What separates a good harness from a fragile one
- Context management: deciding what history, tool results, and instructions to include each turn, so the model has what it needs without being overloaded. This overlaps closely with context engineering.
- Error handling: when a tool fails or the model produces an unusable response, the harness needs a clear fallback: retry, ask the model to fix its output, or stop and report the problem.
- Stopping conditions: a harness needs to know when to stop looping. Without limits, an agent can loop indefinitely, repeat the same failed action, or run up cost.
- Observability: logging each step — what was sent, what came back, what tool ran — makes it possible to debug why an agent did something unexpected.
- Safety boundaries: enforcing which tools can run without confirmation, and which need a human to approve first.
Two systems using the identical underlying model can behave very differently depending on harness quality. A weak harness lets small model mistakes snowball into a bad outcome. A strong harness catches problems early.
An example of harness responsibility
Imagine an agent that fixes failing tests. The model might decide to “delete the failing test” instead of fixing the underlying bug — a plausible but unwanted shortcut. A well-designed harness can:
- Restrict which files or actions the model is allowed to touch.
- Require a human to approve any deletion.
- Detect that a test count dropped and flag it before finalizing.
The model’s imperfect judgment does not become the whole system’s judgment, because the harness adds checks around it.
Isolation and security
A harness is also the security boundary around the model. It should authenticate the user, authorize every action as that user, and enforce per-user permissions before any tool runs. The model can propose an action, but it should not be trusted to decide whether the user is allowed to perform it.
Good harnesses keep secrets out of the context window. API keys, database credentials, and deployment tokens should stay in secure runtime storage, not in prompts or tool results. If code or shell commands can run, use sandboxing such as containers, restricted file systems, limited environment variables, and narrow network access.
Inputs are not automatically safe. Repository files, web pages, support tickets, tool descriptions, and tool results can all contain prompt injection. The harness should treat those observations as untrusted data and pair prompt engineering defenses with allowlists, permissions, and validation. When external tools come through MCP, the harness should still review what each server can reach.
Skills can also affect safety. A skill may include instructions, scripts, templates, and validation steps. The harness decides which skills are installed, when their instructions enter context, and whether any supporting scripts are allowed to run.
Reliability and recovery
Production harnesses need boring software controls:
- Timeouts for model calls and tools.
- Retries with backoff for safe transient failures.
- Idempotent write actions so a retry does not create duplicate tickets, payments, or messages.
- Cancellation when the user stops the task or a deadline is reached.
- Streaming so users can see progress before a long run finishes.
- Crash recovery from checkpointed state and history.
Checkpointing is especially important for long-running agents. If the process restarts, the harness should know which steps finished, which tool results were already observed, and whether any action is waiting for approval.
Observability and evaluation
Harnesses should record traces of model calls, tool calls, decisions, errors, and stop reasons. Sensitive data should be redacted before traces are stored or shared. Cost and latency belong in the same view, because a correct agent that takes too many steps may still be unusable.
Saved traces are useful for evaluation. When a team changes a prompt, model, tool schema, MCP server, or skill, it can replay representative tasks and compare outcomes. This turns agent changes into regression tests instead of guesswork.
Harness engineering vs prompt engineering
Prompt engineering shapes what you ask the model to do in a single call. Harness engineering shapes how the whole system behaves across many calls: what it remembers, what it’s allowed to do, when it stops, and how it recovers from mistakes. A great prompt inside a weak harness can still produce an unreliable system; a solid harness makes the whole agent easier to trust.
The key idea
A harness is the engineering layer that turns a model into a working, bounded agent — managing the loop, the tools, the state, and the limits. Harness engineering is often what determines whether an agentic system is reliable and safe, more so than the prompt or even the underlying model.