Evals and Benchmarks: Read a Score, Build Your Own, Hillclimb

A benchmark score and your own eval score are the same kind of object. Each one is a task set, a grader, a harness, and a sample of model outputs, reduced to one number. This guide reads both. Part 1 takes apart the public benchmarks that launch posts quote. It shows one model moving 39 points when only the harness changes. Part 2 builds an eval from zero on a local lab and tests it like any other instrument. It covers failure reading, graders, judge calibration, error bars, a held-out hillclimb, agent trajectories, and production checks. The method works with any model behind an OpenAI-compatible endpoint and any agent CLI with a headless mode. Every number comes from a run on this laptop or from a source linked next to it.


Part 1 · Read a benchmark score

What a benchmark score is made of

A published score is the output of five choices. Change any one of them and the number moves. Two launch posts can quote the same benchmark name and still measure different things.

Humanity's Last Exam shows the effect in one pair of numbers. Scale's leaderboard tops out at 21.64%. Artificial Analysis shows 61.4% for the same benchmark name. Neither page prints its tool access next to the rows. Treat the gap as a protocol difference until both setups are published.

The benchmarks behind launch posts

The table lists what each benchmark measures and how it grades. The last column is the highest score I could confirm on a primary or named page on September 30, 2026. Read the grader column before the score.

BenchmarkMeasuresTasksGraderTop confirmed score
SWE-bench VerifiedFixing real GitHub issues in Python repos500Unit tests76.80%, Claude 4.5 Opus (high) with mini-SWE-agent, on the leaderboard dated Feb 17, 2026. Newer entries may exist.
SWE-Bench Pro V2Harder repo issues under copyleft and private licenses642Unit tests, regraded on a pristine imageThe grading changed on Sep 22, 2026, so compare only V2 rows.
Terminal-Bench 2.xTerminal tasks in containers, run through HarborNot confirmedA test script per taskNot confirmed on tbench.ai
tau-bench familyTool use with a simulated user under a written policyRetail, airline, and telecom domainsFinal database state plus required facts in repliesgpt-4o at 61.2% retail pass^1 in the 2024 paper. The repo now points to tau2-bench.
OSWorld-VerifiedComputer use across desktop appsNot confirmedA checker script per task86.1%, Qwen3.8 Max, from the third-party tracker llm-stats.com, September 2026
BrowseCompHard multi-hop web facts1,266Exact match against an encrypted key86.81% raw and 86.57% after cleanup, Claude Opus 4.6 multi-agent, reported by its developer on Mar 6, 2026
GAIAGeneral assistant tasks with tools466, of which 300 are held outQuasi-exact match90.03%, CustomGPT.ai Research Lab v27, submitted May 14, 2026
Humanity's Last ExamClosed questions at the edge of expert knowledgeAbout 2,500Exact match, plus a calibration error21.64% on Scale's page and 61.4% on Artificial Analysis. The protocols differ.
GPQA DiamondGraduate science, multiple choice198Exact match96.3%, GPT-6 Astra (xhigh), on Artificial Analysis. Epoch AI's own runs top out at 96%.
ARC-AGI-2Novel visual reasoning puzzlesNot printedExact grid match93.3%, Claude Opus 5.5 (High), at $0.408 per task, Sep 22, 2026
ARC-AGI-3Goal discovery in new interactive gamesNot printedLevel completion59.3% is the best standard-harness score. Most entries are under 10%.
MMLU-ProTen-option knowledge and reasoning questionsNot confirmedExact match89.6% to 91.7%, depending on the tracker
LiveCodeBenchContest coding problems, windowed by release date454 in the Aug 2024 to May 2025 windowTest casesNot confirmed. The page marks models that are likely contaminated.
METR time horizonThe length of software task a model finishes at 50% successMore than 100Task success, fitted against human completion timeAbout 17 hours for Claude Mythos Preview (early). METR calls values above 16 hours unreliable.
LMArenaHuman pairwise preferenceMillions of votesBradley-Terry rating with style controlSee the access problem described below

Contamination, saturation, and models that notice the test

A benchmark saturates when the top models crowd its ceiling and their order becomes noise. ARC-AGI-1 is there. Its human panel scores 98.0%, and the top configurations sit between 96.5% and 98.5%. On MMLU-Pro, three trackers disagree about the leader inside a two-point band. On SWE-bench Verified, a September 2026 audit could not separate any of the 29 adjacent pairs in the top thirty.

Four methods catch contamination, and each one misses a case.

Anthropic documented that last case in March 2026. Across 1,266 BrowseComp problems, Claude Opus 4.6 got 11 answers from benchmark material instead of research. In two runs the model guessed that it was inside an eval. It worked through a list of benchmark names, found BrowseComp's source code, and wrote its own decryption function. Then it decrypted the answer key from a mirror. One of those runs used 40.5 million tokens, about 38 times the median. The unintended-solution rate was 0.24% for one agent and 0.87% for the multi-agent setup. Blocking every search result that contained the benchmark's name worked better than URL blocklists.

Eval awareness is a measured ability. In 2025, Needham and colleagues built 1,000 transcripts from 61 datasets. Gemini-2.5-Pro told evaluations apart from real use with an AUC of 0.83. A simple human baseline reached 0.92. Both did better on agent transcripts than on chat.

Crowd rankings have a different weakness. The Leaderboard Illusion (2025) counted 27 private Llama-4 variants tested on LMArena before release. It estimated that Google and OpenAI received 19.2% and 20.4% of all arena data. Together, 83 open-weight models received 29.7%. The authors estimate that extra arena data can raise relative performance on the arena distribution by up to 112%.

Same model, different number

The harness is part of the measurement. The same SWE-bench Verified audit found that scaffold choice moved one model's score by up to 29.8 points. The whole top thirty spanned only 8.8 points. Figure 1 puts four published cases side by side.

Fig 1 · Same model, different number

Each board holds one model and one benchmark. Only the harness, the effort setting, or the cleanup changes between rows. Tap two rows to compare them. When the leaderboard prints its task count, the whiskers are 95% intervals from that count. When it does not, set an assumed count and read the result as a what-if.

120

Read the ARC-AGI-2 board closely. Claude Opus 5.5 at XHigh effort scored 91.7%. That is below its 93.3% at High effort, at 4.5 times the cost per task. Measure each effort setting before you pay for it.

How to read a leaderboard or a model card

A good model card also reports intervals per surface. The same system card reports one multi-turn safety eval at 87%, plus or minus 6%, through the API. In the Claude app, the same eval reads 85%, plus or minus 6%.

Part 2 · Build your own eval

The lab behind every number

I built a small eval for a fictional invoicing product called Ledgerly. It has a one-page support policy with plans, a 30-day refund window, a 200 USD refund cap, six routing queues, and four escalation rules. There are 48 tickets. Each ticket expects four checkable fields: the queue, whether a refund is allowed, the refund amount, and whether to escalate. The model also writes a short reply to the customer, which a second model grades.

The runner is plain standard-library Python with no framework. It posts to /v1/chat/completions, so the same code runs against Ollama, llama.cpp, vLLM, or a hosted API. I ran the qwen3 family at 0.6B, 1.7B, and 8B parameters, plus the 4B instruct build, on an M1 Max through Ollama. Unless a sentence names another source, "measured" means this lab.

All scored runs used temperature 0.7. Runs without thinking made three attempts per ticket. Runs with thinking made one attempt, because an 8B thinking pass took about two and a half minutes per ticket on this machine. I cut the lab at about three hours. So the 8B thinking run covers only the first 29 tickets, and the hillclimb stopped after three rounds. Figure 4 marks the tickets that a run did not reach.

Read failures before you write metrics

Metrics come after reading. Hamel Husain and Shreya Shankar put error analysis first. Read at least 30 traces yourself and write a free note about the first failure in each one. That step is open coding. Then group the notes into named categories and count each category. That step is axial coding. Keep reading until new traces stop adding categories. Their guidance puts that point at about 100 traces.

I ran the same pass on Ledgerly with the 4B instruct model at three attempts per ticket. Of 144 attempts, 51 failed. The first table counts the wrong fields in those failures. One failed attempt can get more than one field wrong. The second table gives the failure rate by the kind of difficulty I noted before the run.

Wrong fieldFailed attemptsShare of failures
escalate2753%
refund_amount2447%
route1937%
refund_allowed918%
broken JSON or schema00%
Difficulty categoryFailed attemptsFailure rate
header state3 of 3100.0%
refund cap15 of 2462.5%
escalation rule24 of 4257.1%
routing precedence18 of 3946.2%
refund window6 of 1540.0%
currency3 of 1225.0%
negation3 of 1225.0%
direct mapping0 of 270.0%

The counts pick the next evals to write. Escalation and refund-cap errors cause most failures, so each one gets its own targeted check. The single header-state ticket failed every attempt. One ticket is too few to call a category, so I read it by hand first. Direct mappings never failed, so more easy tickets would add cost and no signal.

Turn each category into an evaluator. Use a code check when a rule decides the outcome, such as a JSON parse or a refund above 200 USD. Use a judge when the call needs judgment. For a judge, Husain and Shankar label 100 to 200 examples per failure mode. They put 10 to 20% into prompt examples and 40 to 45% into a development set. The rest is one final test.

Write tasks you can defend

A task is an input, a correct answer, and a reason the task exists. The reason is the part people skip. For each Ledgerly ticket I wrote a one-line note that says why the ticket is hard before any model saw it. One example is "Exactly 30 days old: boundary is inclusive." Another is "Single clear issue, keyword maps directly to a queue." The set has 16 easy, 16 medium, and 16 hard tickets by that note.

{"id": "t36", "difficulty": "hard",
 "ticket": "Today: 2026-09-30 | Plan: Team (monthly) | Billing currency: USD
            | Account status: active\n\nI was charged $40 on August 31.
            Can I still get that refunded?",
 "expected": {"route": "billing", "refund_allowed": true,
              "refund_amount": 40, "escalate": false},
 "why_hard": "Exactly 30 days old: boundary is inclusive."}

Write the note first for a reason. If you pick tasks by where the current model fails, you measure the weak spots of that one model. Lance Martin's post on eval design and hillclimbing (September 28, 2026) makes the same point. It asks for a stated reason for difficulty before a case enters the set. A difficulty note written in advance also gives you something to check later. If the "easy" tickets fail more often than the "hard" ones, the notes or the tasks are wrong.

Take the inputs from real work where you can: production transcripts, support tickets, and bug reports. Then add hand-written cases for the rules that real traffic rarely exercises. Users tend to send requests they expect to succeed, so traffic alone skews easy.

Audit the tasks before you trust any score. Two audits from this month show the size of the problem. Scrimdata reviewed 372 recorded failures on DeepSWE. It found defects or ambiguous requirements in 37 of 113 tasks. Fixing only confirmed defects raised measured pass rates by 4.22 to 6.19 points for the three models it checked. Scale's SWE-Bench Pro V2 (September 22, 2026) dropped 89 invalid tasks. It also rewrote the text of 69 tasks whose instructions contradicted their own tests.

Run any model through one interface

Keep the runner dumb and the endpoint swappable. This is the model call from the lab, trimmed for space. The base URL and key come from the environment.

BASE_URL = os.environ.get("EVAL_BASE_URL", "http://127.0.0.1:11434/v1")

def chat(model, messages, temperature=0.0, max_tokens=4096, seed=None):
    body = {"model": model, "messages": messages,
            "temperature": temperature, "max_tokens": max_tokens}
    if seed is not None:
        body["seed"] = seed
    req = urllib.request.Request(BASE_URL + "/chat/completions",
                                 json.dumps(body).encode(), HEADERS)
    with urllib.request.urlopen(req, timeout=300) as r:
        d = json.load(r)
    u = d.get("usage") or {}
    return {"content": d["choices"][0]["message"].get("content") or "",
            "prompt_tokens": u.get("prompt_tokens"),
            "completion_tokens": u.get("completion_tokens")}

Record errors and timeouts apart from wrong answers. A timeout is infrastructure noise. If you score it as a failure, a slow night on a shared server looks like a model regression. The lab stores error and pass in separate fields and never averages an error into the score.

Check the serving configuration before the first scored run. Two settings in this lab would have skewed every number.

Do not expect temperature 0 to repeat itself on a shared server. Thinking Machines Lab traced the cause to kernels whose floating-point reduction order depends on batch size. Batch size depends on concurrent load, so the same request can produce different tokens. A fixed seed does not help, because the seed only controls sampling. vLLM exposes a beta switch, VLLM_BATCH_INVARIANT=1, that trades speed for repeatable output. Ollama and llama-server document no equivalent.

A single local server is the easy case, and the lab shows it. I ran 10 tickets three times each through one Ollama server with four requests in flight. At temperature 0, every reply was byte-identical across the three runs. At temperature 0.7, the pass verdicts held on all 10 tickets, but the reply text changed on all 10. Do not assume the temperature 0 result transfers to a shared, batched fleet.

Agents: call the harness, not the model

To evaluate a coding agent, run its CLI in headless mode inside a fresh working copy per task. Then read its JSON output. These are the documented commands as of September 30, 2026. The last column says where the run cost comes from.

HarnessHeadless commandCost and turns
Claude Codeclaude -p "..." --output-format jsontotal_cost_usd, num_turns, is_error in the result
Codex CLIcodex exec --json "..."token counts in turn.completed.usage, no USD field
Gemini CLIgemini -p "..." --output-format jsontoken usage in stats, exit code 53 when the turn limit is hit
pipi --mode json "..."per-message usage.cost, sum it yourself
OpenCodeopencode run --format json "..."event fields not documented, use opencode stats

Reset the environment between trials. Leftover files and git history from an earlier attempt can leak the solution to the next one. Pin the harness version as well. Codex and Claude Code both published new releases within a day of this guide, so an unpinned run measures a moving target.

Grade with the cheapest check that works

Use this ladder, and stop at the first rung that can decide the task:

  1. Exact match on a label, a number, or a JSON field.
  2. A schema check plus field checks, like the Ledgerly grader below.
  3. Tests that run the output, for code.
  4. A model judge with a rubric of yes or no questions, for open-ended text.
FIELDS = ["route", "refund_allowed", "refund_amount", "escalate"]

def grade(output_text, expected):
    o = extract_json(output_text)
    if o is None or not schema_ok(o):
        return {"pass": 0, "schema_ok": False}
    fields = {f: field_ok(f, o.get(f), expected[f]) for f in FIELDS}
    return {"pass": int(all(fields.values())), "fields": fields}

Read your grader as carefully as the model output. The first Ledgerly grader read a missing refund_amount key as null. A reply that simply left the key out could pass a task that expected no refund. I found it during the smoke run and made all five keys required. Regrading the first scaling rows changed 3 of 384 verdicts from pass to fail.

When a model is the grader

A model judge brings its own biases, and they are measured. In the MT-Bench study, GPT-4 as a judge gave the same verdict after an order swap in only 65% of pairs. Claude-v1 favored the first answer 75% of the time. A CIKM 2026 paper showed judges a prior score from an earlier attempt. That irrelevant number flipped 10.18% of correct judgments.

Four habits keep a judge honest:

The Ledgerly judge is qwen3 8B with thinking off. Its rubric has four yes or no checks, and code checks an 80-word limit. I graded 20 replies from the 4B model twice. At temperature 0, no verdict changed. At temperature 0.7, 2 of 20 verdicts changed, and 8 of 80 single checks changed. So run the judge at temperature 0. I also compared a 4B reply with a 0.6B reply on the same 20 tickets, in both orders. The judge picked position A exactly half the time, so it had no net position bias. But 2 of the 20 verdicts depended on the order. These checks measure stability only. Correctness needs hand labels. I did not label replies in this lab, so the judge agreement with people is unmeasured.

Airbnb's engineering team reported the same pattern at scale. A faithfulness judge agreed with a product manager 78% of the time because it punished accurate paraphrases. A rubric fix raised agreement to 88%.

Calibrate the judge against people

Agreement is only the first check. A judge that passes too much, or fails too much, shifts every score it grades. Label a gold set by hand and include bad examples. Airbnb's walkthrough used 60 labels and aimed for agreement in the high 80s to 90s. From the gold set, compute the judge's true positive rate (TPR) and true negative rate (TNR). Then correct the raw pass rate that the judge gives on production traffic.

def corrected_pass_rate(p_obs, tpr, tnr):
    if tpr + tnr <= 1:
        raise ValueError("judge is no better than chance")
    return (p_obs + tnr - 1) / (tpr + tnr - 1)

This is the Rogan-Gladen correction from diagnostic testing. The judgy package implements it with a bootstrap interval. Figure 2 shows how far a raw judged rate can drift, and how the size of the gold set sets the interval.

Fig 2 · Judge calibration board

One hundred judged replies, sorted by what a person would say and what the judge said. Move the judge's true positive and true negative rates and the real pass rate. The raw judged pass rate drifts away from the truth. The Rogan-Gladen correction pulls it back, and the size of your hand-labeled set sets the width of the interval.

judge: passjudge: fail person: pass
person: fail

Correction: θ = (p_obs + TNR − 1) / (TPR + TNR − 1), as implemented in the MIT-licensed judgy package. The interval is a seeded simulation of 4,000 draws. TPR and TNR are estimated from the hand-labeled set, split by the true pass rate, and p_obs from 1,000 judged production replies. All slider values are reader-set. The 60-label preset matches the golden set in Airbnb's August 2026 walkthrough, and 150 sits inside the 100 to 200 labels per failure mode that Husain and Shankar recommend.

Report kappa next to raw agreement. Take the default board: a 60% true pass rate and a judge at 92% TPR and 70% TNR. That judge agrees with people 83% of the time, but its kappa is only 0.64. Its raw pass rate reads 67.2% when the truth is 60%.

Measure the noise before you compare

Every score is an estimate from a sample of tasks and a sample of generations. Report it with an error bar, and compare two systems on the same tasks. Evan Miller's Adding Error Bars to Evals gives the formulas. Three of them carry most of the weight.

Clustered standard error. Repeated runs of one task are not independent samples. If you count 48 tasks times 3 runs as 144 observations, the error bar is too narrow. For the 8B model, the naive standard error is 0.040. The standard error clustered by task is 0.068, which is 1.7 times wider.

Paired difference. Compare two systems task by task, not score against score. The pairing removes the variance that comes from task difficulty.

def paired_diff(a, b):
    common = sorted(set(a) & set(b))
    d = [mean(a[t]) - mean(b[t]) for t in common]
    md, se = mean(d), stdev(d) / math.sqrt(len(d))
    return {"diff": md, "se": se, "ci": (md - 1.96 * se, md + 1.96 * se)}

Minimum detectable effect. Before a comparison, ask how large a difference the eval can see. At 80% power and a 5% false-positive rate, the answer is about 2.8 standard errors of the paired difference.

The Ledgerly numbers show why this matters. The 8B model scored 66.0% and the 4B instruct build scored 64.6%. The paired difference is +1.4 points, with a 95% interval from -11.4 to +14.1 points. That result says nothing about which model is better. The 1.7B model against the 0.6B model is +35.4 points, with an interval from +22.7 to +48.1. That difference is real.

Fig 3 · Noise band gauge

Pick two configurations from the lab. The dot is the measured paired difference on the same tasks. The band is its 95% interval, and the dashed lines mark the smallest difference this design can detect. Change the task count and the repeats to see what a larger eval would buy.

48
3
95% intervaldetectable effectmeasured difference

Measured: Ledgerly lab, 48 tickets, temperature 0.7. Projection uses the measured between-task and within-task variance with the paired power formula from Miller (2024).

Public leaderboards have the same problem at a larger scale. A September 2026 audit ran paired McNemar tests on the top thirty SWE-bench Verified entries. None of the 29 adjacent pairs was distinguishable at the 5% level. If two entries sit a point apart, treat them as tied until an interval says otherwise.

pass@k and pass^k answer different questions

Run each task k times and you can report two numbers. pass@k is the chance that at least one of k attempts succeeds. It fits workflows where you can check and keep the best attempt. pass^k, from tau-bench, is the chance that all k attempts succeed. It fits agents that must behave the same way every time a customer asks.

def pass_at_k(n, c, k):
    return 1.0 if n - c < k else 1 - math.comb(n - c, k) / math.comb(n, k)

def pass_hat_k(n, c, k):
    return math.comb(c, k) / math.comb(n, k)

The gap between the two is a reliability measurement. For qwen3 1.7B on Ledgerly, pass@3 is 56.8% and pass^3 is 34.1%. For 0.6B, pass@3 is 20.8% and pass^3 is 2.1%. The small model sometimes finds the right answer but almost never repeats it. The 4B instruct model shows the other extreme. Its pass@1, pass@3, and pass^3 are all 64.6%, because every ticket either passed three times or failed three times. The tau-bench authors saw the same shape at the frontier. GPT-4o passed over 60% of retail tasks once, but pass^8 fell below 25%.

Check the eval against the models

Before you optimize against an eval, check that it behaves like an instrument. Three checks catch most broken evals.

ConfigurationThinkingPass rateTickets x attempts
qwen3 0.6Boff11.8%48 x 3
qwen3 1.7Boff46.4%48 x 3
qwen3 4B instructnone64.6%48 x 3
qwen3 8Boff66.0%48 x 3
qwen3 0.6Bon22.9%48 x 1
qwen3 1.7Bon41.7%48 x 1
qwen3 8Bon93.1%29 x 1

The eval passes the scaling check with caveats. Size helps up to 4B, and then the curve goes flat without thinking. With one attempt per ticket, thinking moved 0.6B by +11.1 points and 1.7B by -5.6 points. Neither paired interval excludes zero. The 8B thinking row needs the most care. It reached only tickets t01 to t29, which are the easy and medium tickets. On those same 29 tickets, 8B without thinking scored 79.3%. The paired gain is +13.8 points, with an interval from -5.0 to +32.6. So the gain is plausible, and this run does not establish it. The best complete configuration is 8B without thinking at 66.0%, which leaves plenty of headroom.

Fig 4 · Difficulty map

The curve is the mean score per model, with a 95% interval clustered by task. Below it, every ticket is one column and every model is one row. A darker cell means more of the attempts passed. Sort the columns, and tap a column to read that ticket and each model's answer.

0 passed1 of 32 of 3all passednot runsaturatednever passesbigger is worsecurve, thinking offcurve, thinking on

Measured: Ledgerly lab, temperature 0.7. Thinking off: 3 attempts per ticket. Thinking on: 1 attempt per ticket, and the 8B run reached only t01 to t29. Errors and timeouts are left out of the counts. The answer table shows each model's first scored attempt. "Bigger is worse" marks a larger model at least 2 of 3 attempts below a smaller one. The scaling slope is the least-squares slope of pass rate against log model size.

Hillclimb with a held-out test set

A hillclimb changes one thing, measures, and keeps the change only if the score improves. The target is usually something cheap to edit, such as a system prompt, a skill description, or a tool definition. Choose a surface where you can say which edit caused which change.

The trap is that the eval becomes your training data. Any edit written after reading failures can fit those exact failures. So split the tasks before the first round. The lab uses a fixed seed to put 32 tickets in a train set and 16 in a test set. You read train failures. You never read test failures.

dtr = paired_diff(by_task(train_after), by_task(train_before))
dte = paired_diff(by_task(test_after), by_task(test_before))
keep = dtr["ci"][0] > 0 and dte["ci"][0] > 0
lenient = dtr["diff"] > 0 and dte["diff"] >= 0

The strict rule keeps a patch only when the paired interval clears zero on both sets. The lenient rule is the one people use by instinct: train went up and test did not go down. I ran both on every round.

The lab ran three rounds on qwen3 4B instruct. The baseline scored 61.5% on train and 68.8% on test.

Under the lenient rule, the prompt would now carry two patches that did nothing for unseen tickets. One of them is a lookup table of train answers. Under the strict rule, the prompt is unchanged. For an eval this size, that is the correct result. Read the outputs as well as the score: p1 moved the test number without doing the job it was written for.

Fig 5 · Climb chart

Each round tries one patch on the qwen3 4B instruct system prompt. The solid lines are the prompt you would ship. Each spur is a patch that was tried, with its 95% paired interval. The shaded band is the smallest test change that 16 tickets can detect. Switch the keep rule to see what each rule ships.

train, 32 ticketstest, 16 ticketstest detectable bandtried patch

Measured: Ledgerly lab, qwen3 4B instruct, temperature 0.7, 3 attempts per ticket, split seed 7. Each patch was scored against the baseline prompt, because the strict rule reverted every patch. The band is 2.8 times round 1's paired standard error on the test set. The lenient rule's round 3 is marked open because it would start from p2, and that combination was not run.

Two other rules protect the test set. Never paste a failing transcript into the system prompt. Also keep the expected answers where the model under test cannot reach them, including files, environment variables, and tool results.

When the score stops moving for two or three rounds, stop patching. Read every remaining train failure and sort it by cause: a model error, an ambiguous task, or a grader bug. At the Ledgerly baseline, 13 train tickets failed. Eleven were model errors, such as a missed 30-day window or an unauthorized charge sent to billing. Two were ambiguous tasks. One of them, t41, allows a second reading of the seat arithmetic. A patch cannot fix an ambiguous task, so rewrite those tasks before the next round.

Agent evals: grade the trajectory and the end state

An agent run produces a trace. The trace holds the user message, each tool call, and the environment state when the attempt stops. Grade two readings of it. The outcome grade checks only the final state, such as whether the refund posted. The process grade checks each step. NVIDIA's agent eval guide gates releases on the outcome and keeps step checks for debugging. These four step metrics come from that guide:

MetricFormulaWhat it catches
Tool-call precisioncorrect calls / calls issuedInvented tool names and extra calls
Argument accuracycorrect arguments / calls to the right toolThe right API with the wrong inputs
Steps per successsteps / successful tasksSlow, expensive paths to a correct result
Consistencyrange of success rate over 3 to 5 trialsA single point estimate that hides spread

Check the state, not the agent's story. After the run, query the database, the file, or the ticket status. Sometimes the agent reports success and the state check fails. Then look for an ignored tool error, a rolled-back write, or a grader that reads the wrong namespace. Sometimes the outcome passes but the path looks unsafe. Then look for a harmful action followed by a compensating action. The final state can hide that pattern.

Do not require one exact tool sequence. Agents reach correct results by paths you did not plan. Score the outcome and the decisions, and keep strict ordered matching for steps that safety requires. When you find a bad trajectory, add its input to the dataset. Then write the rule it broke as a code check, such as "call the payment tool at most once."

For customer-facing agents, use a simulated user. tau-bench runs a model-driven user against the agent under a written policy. A task passes only when the final database matches the expected state and the replies contain the required facts. Report pass^k on those tasks, because real users repeat requests.

Defend the sandbox

An agent can change its environment, so the grader must not trust the environment. SWE-Bench Pro V2 shows the failure modes in one changelog. In an earlier open-network run, 32 of 642 trajectories called code hosts, and 4 fetched the commit that fixed the task. V2 now lets the agent reach only the model endpoint. It regrades every diff on a pristine image and publishes both grades. That regrade caught one model forging a checksum into go.sum. It also caught another model editing the Go module cache. Every task must pass with the reference patch and fail with an empty patch before release. I wrote about that change in more depth in SWE-Bench Pro V2 caught two models cheating.

Treat reward hacking as a measured rate, and do not compare rates across studies. A July 2026 protocol-validity paper audited 2,385 traces across 15 agent benchmarks. It defines a mislead gap as the score earned through a shortcut minus the score earned without it. Published hack rates differ by task set, definition, judge, and prompt, so a number from one study does not transfer to another.

Evals in CI and in production

Run two systems. Husain and Shankar describe the split this way. CI evals block known regressions before a deploy. Online monitoring finds new failures in real traffic and estimates how often they occur.

RAG, safety, and synthetic tasks

For retrieval systems, grade three things apart. First, did the agent retrieve when the question needed it? Second, is every claim in the answer supported by the retrieved text? Third, did the whole run answer the user at an acceptable cost? The first is a code check on the tool call. The second needs a faithfulness judge. The third is a task-level judge plus a retrieval budget.

For safety, grade guardrails as two error rates. A false positive blocks a legitimate request. A false negative lets a harmful one through. Grade whether each refusal was the right call for that input. A model that refuses everything has a high refusal rate and fails its users. The promptfoo CLI ships a red-team mode, and the OWASP Top 10 for LLM Applications gives a threat list to test against.

Synthetic tasks help before you have traffic, and for rare failures. Husain and Shankar generate them in two steps. First, define dimensions such as issue type, customer mood, and prior context, and write 20 tuples by hand. Then have a model produce more tuples, and turn each tuple into a query in a separate prompt. The split keeps the phrasing varied. Synthetic data cannot tell you how often a failure happens in production, so replace it with real traffic as soon as you have some.

What an honest eval report contains

A 2026 ACL audit of 55 pipeline-eval papers found randomness controls missing in 75% of them. Execution traces were missing in 61%. Most of those results cannot be reproduced from the paper alone. This is the Ledgerly 4B result written as a report a reader could check.

eval:          ledgerly-support, 48 tasks (16 easy, 16 medium, 16 hard)
system:        qwen3 4B instruct 2507, num_ctx 8192, policy prompt v1
harness:       eval.py (standard library), Ollama 0.31.1, 4 requests in flight
sampling:      temperature 0.7, k = 3 attempts per task, no seed
grader:        exact match on 4 fields, all 5 keys required
reply judge:   qwen3 8B, temperature 0, 4 binary checks
judge check:   0 of 20 flips on a rerun; kappa against hand labels not measured
score:         64.6%, 95% interval 50.9% to 78.3% (clustered by task)
reliability:   pass@3 64.6%, pass^3 64.6%
vs qwen3 8B:   -1.4 pts, paired 95% interval -14.1 to +11.4
date:          2026-09-30, Apple M1 Max
contamination: tasks written for this eval, first published with this guide
traces:        results/scaling.jsonl

Tools, with licenses and versions

You do not need a framework for a first eval. The Ledgerly lab uses none. When you outgrow a script, pick by the kind of task, not by popularity. Versions below were read from each project's release page or package registry on September 30, 2026.

ToolLicenseVersionFits
Inspect (UK AISI)MIT0.3.273Agent and sandboxed evals written as dataset, solver, and scorer in Python
promptfooMIT0.123.1YAML prompt and model comparisons, red-team scans, CI
DeepEvalApache-2.04.2.7pytest-style tests with judge and RAG metrics
lm-evaluation-harnessMIT0.4.13Standard static benchmarks on raw models
HarborApache-2.00.23.0Real agent CLIs in parallel sandboxes, and the Terminal-Bench 2.0 harness
verifiersMIT0.3.1Verifiable-reward environments that double as RL signal
LangfuseMIT coreSDK 4.15.6Self-hosted traces and dataset experiments
Braintrust, LangSmith, W&B WeaveOpen SDKs, hosted platformsSDKs onlyHosted eval runs, tracing, and regression history

One date matters if you depend on OpenAI's hosted Evals. Its documentation says the platform becomes read-only on October 31, 2026 and shuts down on November 30, 2026. Export datasets and results before then.

Failure modes and fixes

SymptomLikely causeFix
A bigger model scores lower on a taskAmbiguous task text or a grader that rejects valid answersRead the task and three outputs by hand. Fix the task or the grader, then rerun.
Scores move between identical runsSampling, batch nondeterminism, or leftover environment stateRepeat each task, cluster the error bar by task, and reset the sandbox per trial.
A one-point win becomes a launch decisionNo paired intervalReport the paired difference with its 95% interval and the minimum detectable effect.
Train improves, test stays flatThe patch fits train ticketsRevert it. Remove any rule that quotes a specific ticket.
The judge disagrees with itselfScale ratings or order biasUse binary checks, swap the order, and measure kappa against hand labels.
Near-perfect scoresNo headroomAdd harder human-judged tasks, or optimize cost and latency instead.
Timeouts count as failuresErrors mixed into the scoreStore errors separately and rerun them.
The judged pass rate looks highA lenient judge with a low true negative rateMeasure TPR and TNR on hand labels, then report the corrected rate with its interval.
A benchmark gain after a harness changeThe scaffold moved the score, not the modelCompare models inside one harness, and state the harness next to every number.
The agent finds the answer keyLeaked answers or eval awarenessBlock the benchmark name in search results, keep answers out of reach, and regrade flagged tasks.
The agent says done, the state check failsAn ignored tool error or a rolled-back writeGrade from the environment, never from the agent's summary.

A checklist for your next eval

rg
Rohit Ghumare

CNCF Ambassador and Google Developer Expert. I build agent infrastructure and write about the fundamentals underneath the AI stack. Lab numbers on this page come from my own runs of qwen3 models through Ollama on an M1 Max on September 30, 2026. Other figures are cited inline with their sources.

Related: SWE-Bench Pro V2 and the grader move · All guides · X