Part three of a three-part series on building AI systems you can trust. It picks up the support agent from The Model With No Hands: MCP, after How Do You Know It Got Better?: Evals.
A model that can call tools is still not something you would leave alone with real tickets. Take the support agent from a moment ago: it can look up an order and, in principle, issue a refund. But a single call to a model answers a single question, and a real ticket takes a dozen turns (read the complaint, check the order, weigh the policy, draft the reply) and somewhere in that dozen is a refund you cannot let it hand out unsupervised. Between a model that is able to act and a system you can trust to run on its own, there is a missing piece. The piece has a name, and the name is the harness.
A single exchange with a model is one turn: you send a prompt, it sends back a response, sometimes a request to use a tool. That is the entire transaction, and it ends there. To actually resolve a ticket, something has to take that tool request, run it, hand the result back, and call the model again, and keep doing it until the model says it is finished. That something is the harness. The model supplies the judgement, one step at a time; the harness supplies the continuity and the control that turn a string of separate guesses into a job that gets done.
It is worth being concrete about what the harness is, because it is none of the parts you already have a name for. It is not the model, and not the tools; it is the program sitting between them that owns the loop. The chat window you type into is a harness. Claude Code is a harness. The thing you would write to work through support tickets is a harness. Stripped to its bones, it does five plain things in a circle: send the prompt, read the response, notice a tool request, run it, add the result back in, and then round again.
The first thing that loop buys you is a way to stop. A model told to keep working will sometimes keep working well past the point of use: looping, second-guessing, buffing something that was already finished. The harness is what decides the turn is over: the model has declared itself done, or it has spent its allowance of steps, or it has started visibly repeating itself. The model proposes; the harness disposes. The leash is held by the code, not by the thing on the end of it.
The harness sees every tool request before it runs, and that is precisely what makes the refund safe. The harmless lookups it can let through on sight; the refund it can catch and hold, surface to a person, and refuse to act on until someone says yes. The seam between the model asking and the action happening, the gap the protocol deliberately left open, is where the harness installs a checkpoint and posts a guard on it. Because every action passes through this one place, this one place is where your policy actually lives.
You might object that you could just tell the model, in its instructions, never to refund without asking first, and most of the time it would obey. Most of the time is the entire problem. An instruction in a prompt is a request made to a fallible reader, and a fallible reader can be confused, worn down by a long ticket, or talked round by a clever customer. A harness that will not execute a refund without an approval token is not making a request; it is stating a fact that holds even when the model is wrong. The rule worth relying on is the one that does not depend on the model behaving itself.
A harness can only gate what it can see, and that changes how you hand the model its tools. Give it a single all-purpose tool (run whatever command you like) and every action turns up looking identical, so the harness cannot tell a harmless read from an expensive write and must either wave everything through or stop and ask about everything. Promote the actions that matter into named tools of their own (one for refunds, one for sending a message to a customer) and each arrives as a recognisable, typed request the harness can gate, render, and record on its own terms. The shape you give the tools is what decides how much the harness can enforce.
Because every action runs through it, the harness is also the natural place to write everything down: each tool call, each result, each human approval, in order. For an agent that touches customers and money this is not housekeeping; it is how you answer, three weeks later and possibly to someone official, the question of why order 4471 was refunded at all. The harness is the single chokepoint through which the whole trajectory is visible, which is exactly what makes it the only place a record worth trusting can be kept.
Some steps are safe to run at the same time and some are emphatically not, and the harness is what knows the difference. Three independent lookups can happen at once; two refunds against the same order must never overlap. Watching the requests arrive, the harness can fire the safe ones together and line the dangerous ones up one behind the other. The model does not need to understand your rules about what may run in parallel. It just asks, and the harness is what makes the asking safe.
A long ticket is a growing one, too. Every turn adds to the conversation, old tool outputs pile up, and the model’s limited window starts to crowd with things that stopped mattering several steps ago. The harness decides what the model is allowed to see (what to keep, what to compress into a summary, what to quietly drop) and slips in the fresh company context each turn needs. The model’s sense of the situation is, in the end, simply whatever the harness chose to set in front of it.
Real runs also fail in thoroughly unglamorous ways: a dropped connection, a restart, a machine that just goes away mid-task. A plain chat loses its place when that happens. A harness that has been recording each step as it goes can pick the ticket back up from where it stopped, instead of starting over or, worse, half-finishing the work and walking away. For anything meant to run unattended, the ability to resume is most of the distance between a demo and a system you can actually depend on.
None of which means you should always build your own. For coding at your desk, use the harness someone has already written; for ordinary agents, the frameworks exist and are good. You write your own when the rules are the whole point: a hard stop on spending money without a human, a log you could defend in an audit, a budget of steps, writes that must never collide, a run that survives a crash. The harness is where your non-negotiables stop being hopes and turn into structure. The model is the judgement and the tools are the hands; the harness is everything that makes it safe to let the hands move. It is the body, and the conscience, you build around the brain.