Define success before generating
Choose criteria tied to the task. For an extraction prompt, you may care about source-supported values, valid field types, and correct handling of missing data. For a writing prompt, protected facts and reader comprehension may matter more than formatting.
Separate critical failures from preferences. An invented deadline may make a summary unusable; a slightly wordy sentence may only reduce its polish. Combining both into one average can hide the failure you most need to see.
Build a small but varied test set
Start with several ordinary inputs, an ambiguous input, a missing-data case, and a case that previously failed. This is a practical starting set, not a statistically sufficient benchmark. Expand it as you learn where the workflow breaks.
For meeting actions, include a direct commitment, a suggestion, a conditional task, and a transcript with no owner. Keep some cases out of the prompt’s examples so that you are testing transfer rather than repetition.
Use a rubric with observable checks
| Criterion | Pass condition | Critical? |
|---|---|---|
| Evidence | Every extracted action has support in the source | Yes |
| Unknown values | No owner or deadline is invented | Yes |
| Completeness | All explicit commitments are included | Task-dependent |
| Format | Required fields appear with valid types | Yes for imports |
| Readability | A reviewer can identify the next action | Human review |
A rubric is useful only if two reviewers can apply it similarly. Replace “excellent summary” with a check such as “all three recorded decisions appear, and proposals are not labeled as decisions.”
Compare one meaningful change at a time
Record the prompt version, model, relevant settings, date, test input, and result. Run the old and new prompts on the same cases under comparable conditions. If you change the model and prompt together, you cannot tell which change explains the difference.
Prompt version: Model and settings: Test case ID: Expected properties: Observed failures: Critical failure? yes/no Evidence for the assessment: Change to try next: Regression cases to rerun:
Where outputs vary between runs, repeat important cases. Report the number of observed passes and attempts instead of presenting a tiny sample as a stable success rate. This site does not provide benchmarked claims about the builder’s effectiveness.
Diagnose the failure before adding instructions
A factual error may come from an absent source, not vague wording. A parser failure may need a schema feature and validator. A missed item may require splitting a long document or improving retrieval. Prompt changes are one tool among several.
When a revision helps one case but harms another, preserve both cases in the test set. Do not quietly redefine success after seeing the output. The point of evaluation is to expose trade-offs that intuition alone can miss.
Decide when a prompt is ready to use
Use a release rule appropriate to the consequence. A personal brainstorming prompt can tolerate rough edges. A prompt feeding customer records into a system needs strict validation and a recovery route. Human review should be attached to specific failure risks, not added as a vague final instruction.
Keep the previous working version and the test set. Recheck after a model change, a major source-format change, or a new task requirement. A prompt is maintained as part of a workflow, not perfected once and forgotten.
Further reading
Official documentation for the general techniques discussed here. Worked examples and checklists on this page are original instructional material.