Navigation

Case StudiesExperimentationTools

Appearance

Design and product work

AI evals

Also calledEvals, LLM evaluation, Model evaluation

Systematic tests that measure whether an AI feature gives good answers, run every time the prompt, model or data changes.

Because AI output varies, it cannot be checked by eye once and trusted. Evals turn quality into something measured: automated checks for things that can be verified in code, model-graded checks for judgement calls, and human review to keep both honest.

They are only as good as the definition of a good answer behind them, which is why domain experts have become central to AI product teams. Their judgement is the ground truth the evals are built from.

Common mistake

Optimising for the eval rather than the user. Evals should be tested against real usage regularly, or a product can score well on its own tests and still disappoint people.

See all 95 terms