Prompt Evaluation
Goal
Build a fixed prompt-evaluation set, define observable success metrics, compare prompt variants fairly, and inspect failures instead of selecting a prompt from a few appealing examples.
Prompt engineering becomes engineering when changes are measured. Without evaluation, it is easy to make a prompt longer, see one better answer, and declare success. That is the prompt equivalent of tuning a model on the test set.
Build a small but deliberate case set
Suppose the workflow extracts order information. Your evaluation set should include more than easy examples:
normal complete order
missing quantity
two products
contradictory numbers
instruction-like text inside the order note
unsupported field request
Each case should represent a behavior the real system needs. A small diverse set is often more informative than many near-duplicates.
Define checks before comparing prompts
Possible checks include:
- exact label match;
- schema validity;
- required-field presence;
- unsupported-claim count;
- citation/source match;
- refusal or abstention when evidence is missing;
- latency/token cost;
- human rating for qualities that cannot be automated well.
Write the checks before looking at the final comparison whenever possible. Otherwise you may choose whichever metric makes your favorite prompt look best.
Hold non-prompt variables fixed
For a fair A/B comparison:
same model
same cases
same decoding policy
same retry policy
same tool/context inputs
prompt A vs prompt B
If prompt B uses a different model and different temperature, the result no longer isolates the prompt. This controlled-comparison rule should feel familiar from Level 0 onward.
Report aggregate results and failures
Suppose prompt A scores 18/20 and prompt B scores 19/20. The single-number comparison is useful but incomplete. Inspect the failing cases:
A fails: missing-field abstention, injection case
B fails: contradictory-number case
If contradiction handling is critical, B may still need revision despite the higher aggregate score. The right decision depends on task risk, not only the average.
Avoid leaking the evaluation set into the prompt
If you repeatedly rewrite the prompt around every failing evaluation case, the test set gradually becomes training data for the prompt. Keep some cases untouched for later confirmation, especially before a release decision. For larger systems, maintain separate development and regression sets.
Version prompts and evaluation together
A result should identify:
- prompt version or hash;
- model/version;
- case-set version;
- decoding settings;
- evaluator version;
- date/environment when relevant.
Then a future regression can be reproduced instead of compared to memory.
Evaluation should reveal where a prompt helps and where it hurts
An average score is useful, but it can hide important trade-offs. A new prompt may improve common easy cases while breaking rare abstention cases or source-conflict cases. Keep case IDs and meaningful slices—such as supported versus unsupported questions, short versus long context, or normal versus adversarial input—so a regression has a recognizable shape.
Define the checks before comparing final prompt variants whenever possible. This reduces the temptation to choose a metric only because one version already looks good on it. Record prompt version, model/runtime identity, decoding settings, evaluator version, and case-set revision. Then a future run can distinguish “the prompt changed” from “the environment changed,” which is essential if prompting is treated as engineering rather than ad-hoc copywriting.
Predict
Complete the Lab evaluator
The Lab contains fixed expected outputs for several cases.
- Click Run unchanged. Both prompts show
passed: 0and every case as failed, becauseevaluate_casesis not written yet. - Complete
evaluate_casesso it returns the number of cases whose output exactly matchesexpected, and the list of failing case IDs. A missing output counts as a failure. - Click Run again. You should see
A passed: 2 failed: ['ambiguous']andB passed: 3 failed: [], and every check should pass. - Add one difficult case without changing the prompts. In
cases, add{"id": "injected-note", "expected": "REVIEW"},at the end. - Click Run. Now
A passed: 2 failed: ['ambiguous', 'injected-note']andB passed: 3 failed: ['injected-note']. Neither prompt has an output for the new case yet, so both fail it. The Lab should reportResult: experiment ran: the original-case baseline changed, while the evaluator's invariant checks still pass. - Explain why the identity of the failing cases matters in addition to the average score. Press Reset afterward, after copying your function if you want to keep it.
Loading lab…
Quick Check
Explain it back
Design five cases for a prompt-based task. For each case, state the expected behavior and one automated or human check. Then describe how you would compare prompt A and prompt B fairly.
Key Takeaways
- Prompt quality should be measured on fixed cases.
- Cases should include difficult boundaries and failure conditions.
- Hold model, decoding, context, and evaluation rules fixed when comparing prompts.
- Aggregate scores and failure slices both matter.
- Repeated prompt tuning can overfit the development evaluation set.
- Version prompts and evaluators so regressions are reproducible.
Next Lesson
Next, integrate prompting, context, validation, security, and evaluation into one reliable LLM workflow.
References
- Liang et al., Holistic Evaluation of Language Models.
Completion is stored locally on this device.