Part two of a three-part series on building AI systems you can trust. It follows How Do You Know It Got Better?: Evals and continues with The Body Around the Brain: The Harness.
Picture the assistant your support team uses. It drafts replies that are warmer and clearer than anything the team wrote by hand, it never tires, and it is, in one specific and maddening way, useless: to find out whether order 4471 has actually shipped, a human still has to open the order system, read the status, and paste it back into the chat. The model can reason about the order, draft an apology for the order, recommend a remedy for the order, and cannot see the order. It is a brilliant new hire who has read every manual in the building and been given no login. Closing that gap, between a mind that can think about your systems and a mind that can reach into them, is the whole subject of this piece.
The first idea you need is tool use, and it is smaller than it sounds. You can hand the model a tool (say, one that looks up an order by its number) and tell it the tool exists. When the model decides it needs the order, it does not go and fetch it. It emits a request: look up order 4471. Your code receives that request, runs the actual lookup, and hands the result back. The model decides what it wants; your code decides whether to honour the request and how to carry it out. The line under all of this, the one everything else rests on, is that the model never executes anything itself. It only ever asks.
A tool, described to the model, is three things: a name, a sentence of plain English saying what it does, and a typed list of the inputs it expects. The model chooses among its tools the way you choose a kitchen drawer, by the label on the front. This makes the description not documentation but the interface itself. “Look up the current shipping status of a customer order by its order number” is what makes the model reach for the tool at the right moment; a vague label, and it reaches at the wrong ones, or not at all.
So you wire the order-lookup tool into your support assistant, and it works. Then you want the same capability in the agent that triages tickets overnight. Then a colleague wants it in a script of their own. Each time, you re-implement the same wiring against a different client, in a slightly different shape, and each copy is a thing that can drift from the others. The tool itself is fine. It is simply trapped inside one application, and you are about to build the same small bridge three more times. This is the point where a standard starts to look less like bureaucracy and more like mercy.
That standard is MCP, the Model Context Protocol, and it is the agreement that ends the re-wiring. It is a fixed way for a program to announce “here are the tools I offer, and here is exactly what each one takes.” Write the tool once, behind something that speaks the protocol, and any client that also speaks it can use the tool with no bespoke glue. It is, near enough, a wall socket for capabilities: the shape of the plug is agreed in advance, so the appliance and the outlet no longer need to know anything about each other to work together.
The two halves of that arrangement have names worth keeping straight. The server is the side that exposes capabilities; your order lookup lives behind one. The client is the side that consumes them: the assistant’s harness, whether that is a chat app, your overnight agent, or a colleague’s script. The value of the split is the ignorance it permits. The server does not know or care which model is calling it; the client does not know how the lookup is actually implemented. Either side can be rebuilt without so much as a note to the other, which is the same decoupling that lets you change the plumbing in your house without telling every lamp.
A server is not confined to actions, either. Alongside tools (the verbs) it can expose resources, which are the nouns the model is allowed to consult: your refund policy, say, offered up as a document the model can pull in when it needs the rules rather than guessing them. It can also offer ready-made prompt templates for jobs the team does over and over. The mental model is a small counter with a few labelled buttons and a shelf of reference binders behind it, all described in the same standard language.
How the two halves actually talk is a separate question from what they say, and keeping those separate is deliberate. The same server can run as a little process on your own laptop, spoken to directly, or sit behind a URL for the whole company to share. Moving it from the first to the second does not change a single tool, only the wire underneath. You build the capability once and decide where it lives later.
Now add the second tool, the one that earns the architecture its keep: issuing a refund. Looking up an order is harmless; refunding one moves money, and is not. Here the shape of the thing pays off. Because the model only ever emits intent (refund order 4471) and the server is what actually acts, the gap between the asking and the doing is a place you can stand a human being. The assistant proposes the refund; a person approves it; the server carries it out. The point of consent is not bolted on after the fact as a nervous afterthought. It sits exactly where the protocol already left a seam.
There is a credential hiding inside that refund, and how it is handled is the most important property of the whole arrangement. The server holds the key to your order system. The model does not. The model says refund 4471; the server, which holds the credential, decides whether that is permitted and does the deed. The secret lives behind the protocol, in code you own, and never enters the model’s context or its prompt, which means nothing the model reads, and nothing a customer types into it hoping to be clever, can pry the key loose. The safety is not a rule the model is asked to follow. It is a wall the model is standing on the wrong side of.
Much of this you should not build at all. There are ready-made servers for GitHub, for Slack, for the systems every company has; connect those and spend your effort elsewhere. You build your own server for the thing that is uniquely yours: your internal order system, your warehouse, the bespoke process no vendor has ever heard of. A custom server is the smallest object that converts “the assistant can talk about your business” into “the assistant can act inside it,” and it is worth building precisely when the system it reaches is one only you have.
Once the server exists, the arithmetic changes in your favour. Every client gets the capability with no fresh glue; the credential stays put; the point of consent is fixed in one place instead of scattered across applications; and adding the next tool is one more function rather than another whole integration. The cost is honest and worth saying plainly: you now own a server, which can break, which needs maintaining, and whose tool descriptions have to be good or the model will keep reaching for the wrong drawer at the wrong moment. Most of the time the trade is worth it, because what you have built is not really a feature. It is a socket. And the machine you have wired it to is the one nobody else can reach.