prompt-engineering.si
Evaluation

How to evaluate prompts: a small test set and clear criteria

Prompt evaluation compares the outputs a prompt produces against criteria defined before you see the answer. A single impressive response is a demonstration. Repeated performance on representative cases is more useful evidence.

Define success before generating

Choose criteria tied to the task. For an extraction prompt, you may care about source-supported values, valid field types, and correct handling of missing data. For a writing prompt, protected facts and reader comprehension may matter more than formatting.

Separate critical failures from preferences. An invented deadline may make a summary unusable; a slightly wordy sentence may only reduce its polish. Combining both into one average can hide the failure you most need to see.

Build a small but varied test set

Start with several ordinary inputs, an ambiguous input, a missing-data case, and a case that previously failed. This is a practical starting set, not a statistically sufficient benchmark. Expand it as you learn where the workflow breaks.

For meeting actions, include a direct commitment, a suggestion, a conditional task, and a transcript with no owner. Keep some cases out of the prompt’s examples so that you are testing transfer rather than repetition.

Use a rubric with observable checks

CriterionPass conditionCritical?
EvidenceEvery extracted action has support in the sourceYes
Unknown valuesNo owner or deadline is inventedYes
CompletenessAll explicit commitments are includedTask-dependent
FormatRequired fields appear with valid typesYes for imports
ReadabilityA reviewer can identify the next actionHuman review

A rubric is useful only if two reviewers can apply it similarly. Replace “excellent summary” with a check such as “all three recorded decisions appear, and proposals are not labeled as decisions.”

Compare one meaningful change at a time

Record the prompt version, model, relevant settings, date, test input, and result. Run the old and new prompts on the same cases under comparable conditions. If you change the model and prompt together, you cannot tell which change explains the difference.

Prompt version:
Model and settings:
Test case ID:
Expected properties:
Observed failures:
Critical failure? yes/no
Evidence for the assessment:
Change to try next:
Regression cases to rerun:

Where outputs vary between runs, repeat important cases. Report the number of observed passes and attempts instead of presenting a tiny sample as a stable success rate. This site does not provide benchmarked claims about the builder’s effectiveness.

Diagnose the failure before adding instructions

A factual error may come from an absent source, not vague wording. A parser failure may need a schema feature and validator. A missed item may require splitting a long document or improving retrieval. Prompt changes are one tool among several.

When a revision helps one case but harms another, preserve both cases in the test set. Do not quietly redefine success after seeing the output. The point of evaluation is to expose trade-offs that intuition alone can miss.

Decide when a prompt is ready to use

Use a release rule appropriate to the consequence. A personal brainstorming prompt can tolerate rough edges. A prompt feeding customer records into a system needs strict validation and a recovery route. Human review should be attached to specific failure risks, not added as a vague final instruction.

Keep the previous working version and the test set. Recheck after a model change, a major source-format change, or a new task requirement. A prompt is maintained as part of a workflow, not perfected once and forgotten.

Further reading

Official documentation for the general techniques discussed here. Worked examples and checklists on this page are original instructional material.

Keep going