Design and product work
AI evals
Also calledEvals, LLM evaluation, Model evaluation
Systematic tests that measure whether an AI feature gives good answers, run every time the prompt, model or data changes.
Because AI output varies, it cannot be checked by eye once and trusted. Evals turn quality into something measured: automated checks for things that can be verified in code, model-graded checks for judgement calls, and human review to keep both honest.
They are only as good as the definition of a good answer behind them, which is why domain experts have become central to AI product teams. Their judgement is the ground truth the evals are built from.
Common mistake
Optimising for the eval rather than the user. Evals should be tested against real usage regularly, or a product can score well on its own tests and still disappoint people.