Part one of a three-part series on building AI systems you can trust. It continues with The Model With No Hands: MCP and The Body Around the Brain: The Harness.
Here is a situation that will be familiar to anyone who has built something with a language model. You have an assistant that reads an incoming customer message and does two things: it tags the message as Billing, Technical or Account, and it drafts a reply. It works. You demonstrate it, the room nods, and it goes live. A week later you notice it mislabels a certain kind of refund request, so you adjust the prompt. The refund case is fixed. You check it twice. You ship the change. And here is the question you cannot answer: did everything else stay as good as it was? You changed one instruction in a system that responds to thousands of messages you have never read, and you are about to find out whether you helped or quietly broke something, in production, in front of customers. This is the problem evals exist to solve.
The instinct, when you cannot answer that question, is to eyeball a few outputs and trust your gut. Engineers call this the vibe check, and it works right up until it doesn’t. A vibe check scales to about five examples and no further; the sixth blurs into the fifth, and the failure you needed to catch was in the four hundred cases you didn’t look at. An eval is what you reach for instead. At its simplest, an eval is a test for software whose output you cannot predict exactly: the same idea as the automated tests engineers already write, bent to fit a system that does not give the same answer to the same input twice.
The first thing an eval needs is a fixed set of examples to run against. For the triage assistant, that is a collection of real customer messages (forty of them, say), each paired with the answer you wish the assistant would give: this one is Billing, that one is Technical. The word that matters is fixed. The set does not change between runs. This is the whole trick, and it is easy to miss: the value of the test set is not that it is large or clever but that it is the same every time, because only a frozen ruler lets you compare two measurements. If the examples shifted each run, a change in the score could mean the assistant improved or merely that you had measured a different thing. You freeze the set so that any change in the result can be blamed on the assistant and nothing else.
The second thing an eval needs is a way to turn an output into a score, which is called the grader. For the tagging task the grader is almost insultingly simple: did the assistant’s label match the one you wrote down? Yes is one point, no is zero. That simplicity is deceptive, because choosing the grader is where the real judgement lives. The grader is the part of the eval that decides what “good” actually means. For a label, good means exactly correct. For other tasks, as we are about to see, good is far slipperier, and the grader has to grow teeth.
With a frozen set and a grader you can finally produce the thing you were missing: a number. Run all forty messages through the assistant, grade each label, and the share it got right is your accuracy. Suddenly “I think it got better” becomes “accuracy was eighty-two per cent before my change and eighty-eight after,” which is a sentence you can stand behind. More to the point, it is a sentence that can be challenged in public and survive, because you can show exactly how you arrived at it. The feeling has become a measurement, and measurements can be argued with, reproduced, and trusted.
And then the drafted reply ruins everything. The label had a single correct answer; the reply does not. There is no canonical sentence the assistant was supposed to write, no string to match against, nothing for the exact-match grader to do. Two completely different replies can both be excellent, and a third that shares most of their words can be terrible. This is the point where most people give up on measuring and slink back to the vibe check, and it is exactly the point where evals get interesting, because the problem is not that good cannot be defined; it is that good has to be defined as more than one thing.
The fix is to stop asking whether the reply is correct and start asking whether it has the properties a good reply would have. Did it actually address the question the customer asked? Is the tone right for someone who is already annoyed? Did it invent a policy that does not exist? Each of these is a small, answerable question, and together they form a rubric, a checklist that converts a vague sense of quality into a set of specific things you can check one at a time. A rubric is just the grader for open-ended work, broken into pieces small enough to grade.
But who does the checking? Forty replies against four rubric questions is a hundred and sixty judgements per run, and you will run the eval dozens of times; no human is going to sit through that, and the moment it becomes a chore the eval quietly dies. So you hand the rubric to another language model and ask it to grade (the practice known, plainly enough, as LLM-as-judge). You give the judge the customer message, the assistant’s reply, and the rubric, and it returns a score for each criterion. The critical thing to hold onto is that the judge is itself a fallible system. You have not escaped the original problem; you have moved it. Which is why, before you trust the judge, you check it the way you would check any new instrument: grade a few dozen replies by hand, have the judge grade the same ones, and confirm it agrees with you often enough to rely on. An eval whose judge you never verified is a thermometer you never held to boiling water.
There is a subtler trap inside judging, which is that asking a model to put an absolute score on a single reply is noisy: its sense of what “seven out of ten” means drifts from one reply to the next. It is far steadier at a different question: shown two replies to the same message, which is better? People are the same; we are poor at rating a wine out of a hundred and good at saying which of two glasses we prefer. So when the scores feel unreliable, you switch from grading replies alone to comparing the old assistant’s reply against the new one, message by message, and counting how often the new one wins. You stop measuring quality in the abstract and start measuring the only thing you actually care about: did this change make it better than what it replaced?
That comparison is the engine of the whole exercise, because it turns the eval into something you run on every change rather than once. Edit a prompt, run the eval. A new model is released, run the eval. The reward is the sentence you could never produce before: the refund case is fixed, overall accuracy is up four points, and (here is the part the vibe check would never have caught) the assistant’s recall on urgent messages dropped, so in teaching it to handle refunds you taught it to relax about emergencies. That last clause is the entire reason evals are worth the trouble. They do not just tell you when you succeeded; they tell you what you broke while succeeding, which is the failure that ships to customers precisely because nobody was looking at it.
None of this makes the assistant correct, and it is worth being honest about that. An eval is only as good as the cases inside it, and your forty messages can only ever catch the failures they happen to contain. The first time production surfaces a kind of message your set never imagined (a customer writing in two languages, a refund that is also a complaint), the eval is silent, because the case was not there. The discipline, then, is to treat the dataset as a living thing: every genuinely new failure you find in the wild gets written down and folded back in, so the eval slowly grows to cover the long tail of reality. Evals do not prove your system is right. They stop it from re-breaking in the ways you already know about, which is what tests have always done and exactly as much as you should ask of them.
Like anything worth having, evals cost something: the work of curating the cases, the tokens spent letting a judge read every reply on every run. So you do not eval everything. You eval the part that is load-bearing, the decision that hurts when it is wrong, and you leave the rest to the vibe check it deserves. And you keep one last danger in view: the moment a number becomes the target, people begin gaming it, and an assistant tuned until it aces your forty cases may simply have learned your forty cases. The eval is a yardstick, not the thing being built, and a yardstick you start to worship stops measuring anything. Used well, though, it buys you the one thing the demo never could: the ability to change your system on purpose, and to know, rather than hope, that you left it better than you found it.