Evals
evaluations, benchmarks, testing
In one sentence
A fixed set of test cases, scored the same way every time, that shows whether a change to a prompt, model or harness made your system better or worse.
Why it matters
Output varies from run to run, so three good-looking tries is an anecdote. Evals are unit tests for behaviour: cases collected from real failures, graded by code, by humans, or by another model. A public benchmark tells you about the model in general, never about your use of it. Teams that ship AI features reliably tend to have these before they have fancy prompts.
Sources
Back to all termsReady to make an impact?
Turning complex ideas into clear, user-centered product experiences.