Evals

evaluations, benchmarks, testing
In one sentence

A fixed set of test cases, scored the same way every time, that shows whether a change to a prompt, model or harness made your system better or worse.

Output varies from run to run, so three good-looking tries is an anecdote. Evals are unit tests for behaviour: cases collected from real failures, graded by code, by humans, or by another model. A public benchmark tells you about the model in general, never about your use of it. Teams that ship AI features reliably tend to have these before they have fancy prompts.

Back to all terms

Ready to make an impact?

Turning complex ideas into clear, user-centered product experiences.