Skip to content
IMERGIX — home

AI & automation · June 2026 · 5 min read

Evaluation sets are cheaper than incidents

Forty labelled examples written in an afternoon catch more regressions than a month of prompt tuning.

Prompt tuning without a labelled set feels productive until production disagrees. The cheaper habit is writing forty examples in an afternoon and running them on every change.

Incidents teach the same lessons more slowly and in public. Evaluation sets teach them before customers notice.

Why a small set beats endless tuning

Tuning optimises for the last conversation. An eval set freezes the failure modes you already care about so a 'better' prompt cannot silently erase them.

  1. 01
    Capture known failures firstStart with the cases that already embarrassed you. Perfect coverage comes later; regression coverage comes now.
  2. 02
    Score what operators scoreIf reviewers care about field accuracy, score fields. If they care about tone, score tone. Do not invent metrics nobody uses.
  3. 03
    Run on every changeModel bump, prompt edit, retrieval tweak — same suite. If it is optional, it will be skipped under deadline.
Cost of learning — same failure mode
Writing a 40-example set3h
Prompt tuning without evals20h
One production incident14h
Mint = labelled eval time. Navy = incident and hotfix time.

How small is useful

Forty labelled examples written in an afternoon catch more regressions than a month of prompt tuning.

Forty labelled examples written in an afternoon catch more regressions than a month of prompt tuning.

When the suite should run

On every pull request that touches prompts, tools, or retrieval. Weekly against a larger holdout. After every incident, add the case that escaped.

A short test

Change one line of the system prompt and ask whether anyone would notice before a customer does. If the answer is no, you do not have an evaluation set yet.

Author TBDAI & automation lead · IMERGIX

Building without evals?

Bring three failure cases and we will sketch the first suite with you.

30 minutes · No pitch · A written summary afterwards