AI, usefully · 4 min read ·
You changed the prompt. Did the AI actually get better?
You revise a prompt, run it again, and the answer looks better. It is clearer, more complete, perhaps a little more confident.
Then you try a different question. Something that worked yesterday now fails.
This is why AI development needs evaluations, usually shortened to evals: repeatable tests with defined inputs and criteria for success. They help answer a specific question: did this change improve the system across the work we need it to do?
A validator checks an individual result. An evaluation compares behaviour across a collection of cases. Anthropic’s guide to agent evaluations makes another useful distinction: an agent’s claim that it completed a task must be checked against what actually happened.
For an analytics assistant, that means “analysis completed” is not a passing result. The numbers, interpretation and handling of missing evidence all need to hold up.
Write the answer key before improving the prompt
Suppose an assistant prepares a weekly sales summary. It reads a small dataset, calculates changes and explains which categories need attention.
Before adjusting its instructions, create a few examples where you can independently establish the expected behaviour.
Here is a starting set. These are fictional test cases:
| Input condition | Expected behaviour |
|---|---|
| Sales rise from 100 to 120 | Reports a 20% increase |
| Previous-period sales are zero | Explains that ordinary percentage growth is undefined |
| One day’s data is missing | Flags the incomplete comparison |
| Returns exceed sales | Preserves the negative net-sales value |
| Sales decline without an explanation in the data | Reports the decline without inventing its cause |
The answer key should describe requirements, not prescribe one perfect paragraph. “Sales increased by 20%” and “Sales were up one fifth” can both be correct.
For the missing-day case, decide what the product should do. Should it stop? Compare only complete dates? Present a provisional result? That is a product decision to settle before grading the assistant.
Writing tests often exposes decisions the original specification never made.
Make the comparison fair
Save the current version as your baseline: the system you will compare changes against.
Run both versions on the same questions, dataset and tool responses. If the new version receives cleaner data or a newer policy, you cannot attribute the improvement to the prompt.
For this sales assistant, I would record the prompt version, model version, input snapshot, output and check results. Then change one component first—perhaps the instruction for incomplete periods.
Repeat important cases. Model outputs can vary between attempts, so a single success can overstate reliability. Anthropic recommends multiple trials when evaluating agents for this reason.
Keep some cases aside while developing. If you repeatedly rewrite the prompt around the same five examples, it may become excellent at those examples while remaining fragile elsewhere.
A new category name, different date range or unfamiliar wording can reveal that weakness.
Give each check the right judge
For “100 to 120,” calculate the expected change directly. There is little value in asking another language model to decide whether the arithmetic feels right.
Interpretation needs a different approach. Did the summary turn a possible explanation into a fact? Did it omit a caveat that would change the recommendation?
A human can review those questions using explicit criteria. A model can help scale that review, but its judgments need checking too. Google’s guidance on evaluating a judge model describes comparing model-generated ratings with human ratings to assess whether the judge is suitable for the use case.
For our example, a useful reviewer instruction would be:
Review this sales summary against the supplied data and expected behaviour. Check numerical accuracy, period completeness and unsupported causal claims separately. For each failure, quote the relevant sentence and identify the conflicting evidence. Do not reward confident wording or extra length.
Try the reviewer on an intentionally flawed summary first. If it approves an invented explanation, improve the grading process before relying on its scores.
Read the failures beneath the average
Imagine version A passes 16 of 20 cases. Version B passes 18.
That sounds like progress. But suppose B fixes three formatting problems and introduces one new failure: it confidently reports a decline when a data feed is incomplete.
Whether that trade-off is acceptable depends on the product. The overall score cannot make the decision for you.
Keep results visible by failure type. For this assistant, I would separate calculation errors, missing-data handling, unsupported explanations and readability. Define any failures that block release before seeing the results.
Also inspect cases that changed from passing to failing. Those are regressions—previously working behaviour that the update broke.
A small exercise to start with
Choose one recurring AI task and collect ten permitted, anonymised inputs. Write the expected behaviour for each before running the assistant.
Include straightforward cases, ambiguous ones and at least one where the correct response is to request missing information.
Run your current prompt, make one deliberate change, then run the same set again. Record which cases improved, which regressed and which still fail.
Ten cases will not establish production reliability. They will give you a more useful basis for the next edit than “this answer sounds better.”
Over time, add unfamiliar examples and real failures. Your evaluation set should grow with your understanding of the work—and remain separate from the claim that the work is solved.