The direct answer
An AI evaluation — an eval — is a test for an AI system. You feed the model a set of inputs, apply grading logic to its outputs, and get a score. It's the closest thing the AI world has to the unit tests software engineers run on ordinary code. Anthropic's engineering team defines it plainly: give an AI an input, then apply grading logic to its output to measure success.
Teams run evals every time they change a prompt, swap in a new model, or ship a feature. If the score drops, something regressed — caught before a single user sees it. Evals don't prove a model is "smart"; they prove it still does the specific jobs you tested, as well as it did last week.
How it works
The anatomy of an eval has a few standard parts. A task is a single test with defined inputs and success criteria. Each attempt at a task is a trial — and you run multiple trials, because model outputs vary between runs. A grader is the logic that scores performance, and graders come in three flavors: code-based (exact string match, unit tests, regex checks — fast, objective, brittle), model-based (another AI judging against a rubric — flexible, but needs calibration), and human (expert review — the gold standard, slow and expensive).
The build flow, per OpenAI's guide, is three steps: describe the task as an eval, run it against test inputs, then analyze the results and iterate on the prompt. Run that on a fixed bank of tasks and the scores become baselines and regression alarms — the moment a change breaks something that used to work, the numbers tell you.
A simple example
As an illustration, imagine you run a support desk and want a model to sort incoming tickets into Hardware, Software, or Other. You gather a few hundred real tickets, each labeled by a human. Your grader is simple: does the model's label match the human's label? Run the eval and you might get 87% agreement.
Now you tweak the prompt and re-run. The score moves to 91% — keep the change. It drops to 80% — revert it. One number, one decision, no arguing about vibes. Multiply that loop across dozens of behaviors and you have the basic machinery of AI quality control.
Why it matters
Evals are what let AI products move fast without breaking things. Anthropic's team notes that groups without evals end up "flying blind" — debugging reactively after users complain — while teams with evals can test a new model against hundreds of scenarios and upgrade in days instead of weeks. When a stronger model comes out, the eval suite is the difference between a confident swap and a month of manual testing.
They also force a useful discipline: writing an eval requires you to state what "good" actually means, which is often the first time anyone has stated it precisely. Two engineers can read the same spec and picture different behavior; a shared eval suite settles the argument.
The common misunderstanding
A high eval score is not proof of general competence.
An eval measures performance on the tasks it contains — nothing more. A model can ace a narrowly written test while failing the same task phrased differently, or find a loophole that satisfies the grader while missing the point entirely. Anthropic describes an agent that "failed" a flight-booking eval by discovering a better solution than the grader allowed. Scores are signals about specific, tested behavior, not a certificate of intelligence. Always ask: does this test measure what my users actually need?
What changed recently
The biggest shift is toward agent evals: testing systems that work across many turns — calling tools, changing state, adapting to intermediate results — instead of single question-and-answer checks. That demands grading the conversation transcript and the final outcome separately, and mixing grader types the way a good test suite mixes unit and integration tests.
On the tooling side, the industry is converging on datasets-plus-graders as the durable pattern. OpenAI, for example, is winding down its dedicated Evals platform in late 2026 in favor of dataset-based evaluation — a sign that the test data and the grading logic matter more than any one dashboard.
Try it on PlainLogic
Debug the Build on PlainLogic is a small taste of this mindset: you're handed something broken, and the only way forward is to define what "fixed" looks like and test against it.
Sources
- PRIMARY SOURCEOpenAI: Working with evals
- PRIMARY SOURCEAnthropic: Demystifying evals for AI agents