Harness And Agents

Agent FrameworksAI Agent MemoryContext EngineeringGaurav Dadhich2026-09-283 min read
Harness And Agents
On this page
  1. Two circles and an overlap
  2. Nothing about the harness is new
  3. Who decides what
  4. So, what does a harness look like?
  5. Why the distinction matters

LLM calls are stateless, and a model on its own cannot act. It turns text into more text, then stops. Anything we call an agent is that model placed inside software that calls it in a loop, runs the tools it asks for, feeds the results back, and decides what it is allowed to touch. That software is the harness, and most of the confusion about agents comes from using the two words as if they meant the same thing.

Two circles and an overlap

Blog content image

Picture two circles. One is the model, which thinks and decides what to do next. The other is the harness: the loop that keeps calling the model, the tools it can reach (built-in ones as well as anything connected through MCP), the skills and memory it loads into context, compaction when that context fills up, the permissions that fence it in, and the system prompt that shapes its behavior.

An agent is only the overlap. A model with no harness answers one question and stops; a harness with no model is a car with nobody in the driver's seat. Put them together, give them a task, and the model asks to read a file, the harness reads it, the model picks the next step, and the cycle repeats until the job is done.

Nothing about the harness is new

Every agent has had a harness since the first ReAct-style loop, because somebody had to write the code that calls the model again after a tool returns. We called it scaffolding or the orchestration layer, and LangChain and AutoGPT were harnesses long before anyone used the word.

"Harness" came over from testing, where a test harness is the rig you strap something into so you can run it. It caught on for agents once the harness became the product (Claude Code is the obvious example), and once it became clear how much the harness moves results. When researchers at Queen's University held one model fixed (Qwen3-Next-80B-A3B-Instruct) and ran 35 successive releases of the Qwen Code harness against the same 50 SWE-bench Verified tasks, the share of tasks resolved moved between 23% and 39%, and the later releases used nearly twice the tokens of the early ones (study).

Who decides what

Starting a harness with a model and a task is the main agent; there is no separate step where the harness decides to spin one up. Subagents work differently, and the decision there sits with the model. The harness offers a tool for delegating work, the model decides mid-task that a search is big enough to hand off, and the harness checks the request, starts a fresh loop with its own prompt and a clean context, and passes the result back.

A useful way to hold it: the model is a manager deciding to bring in a contractor, and the harness is the operations team that sets the rules and hands over the finished work. Some harnesses do start loops on their own, for example to summarize a conversation when context runs out, or a fixed reviewer step in a pipeline, and in those cases the harness really is the one deciding.

So, what does a harness look like?

The model sits outside the harness and plugs into it, which is why one harness can swap one model for another and still be the same harness. Skills are not tools: a skill is a playbook the harness loads into context when a task calls for it, while a tool is something the model calls, and MCP is one more source of those tools. The main agent is the whole harness running with a model; only subagents live inside it, as separate loops. Behavior is only partly static, because the config and the system prompt are fixed while how proactive an agent actually is comes from the model following them.

Blog content image

Why the distinction matters

Memory is where this stops being a vocabulary debate for me. Since every model call starts from zero, whatever continuity an agent has across sessions or users was engineered into the harness by somebody, and that is the layer we build Maximem Synap for. Swapping in a stronger model while keeping a weak harness tends to buy less than people expect, because the model can only act on what the harness remembers and allows.

So the one-line version I now give people: the model decides what to do, the harness decides what it is allowed to do and then does it, and the agent is the two of them getting a job done.

From the team at Maximem

Stop rebuilding agent memory from scratch

Maximem Synap is the context management layer we built after hitting every problem in this post ourselves. Persistent recall across sessions, entity resolution and conscious forgetting, in Python, TypeScript and REST.

Related posts