Skip to main content

Verification and Evidence

An agent tells you the feature works. The tests pass. There’s even a screenshot. Then you open the app, and the button does nothing.

Nobody lied, exactly. The agent reported evidence that couldn’t have failed. The trick is telling the difference between a check that measures something and a check that merely agrees with the agent.

The ideal

Start with the oracle: whatever decides that the work is done. It might be a test suite, a linter, a script, or a person. What matters is that the answer comes from somewhere other than the agent’s opinion of its own work.

The ideal is that the oracle isn’t something the agent being judged can reach. That can be as simple as:

  • Tests, linters, or static analysis (tools that inspect code without running it), as long as the agent can’t edit them. Protecting the oracle covers how.
  • Continuous integration (CI), a service that runs your checks on every push, running on a system the agent can’t touch.

It’s not lost on me that this isn’t always possible. In that case, the default option is going to be you. Your review is incredibly valuable, but only sometimes. It’s a lot less valuable when all it produces is “nope, that button still isn’t aligned correctly,” pasted back one round at a time. That’s the meat proxy problem again.

The definition of done

Before the agent starts, write down what “done” means. A definition of done is the evidence half of a task contract: you write it in the task prompt, the spec, or the ticket before the agent starts. Each acceptance criterion in it is one behavior claim plus the check that would disprove it. A definition of done names:

  • A behavior claim.
  • A command, inspection, or observation that could disprove it.
  • What a failure looks like.
  • What evidence to keep.
  • Who owns changing the check.

That last item matters more than it looks. Whoever can change the check can make it pass.

Three ways evidence lies

  • The false green: ./check.sh 2>&1 | head exits 0 even when check.sh exits 7, because a pipeline reports the exit status of its last command, and that’s head. In bash, echo ${PIPESTATUS[@]} prints 7 0: the truth and the lie, side by side. No model required.
  • The test that tests nothing: Give an agent a buggy slugify function and three visible tests, and a special-cased solution (one that hard-codes the answers to exactly those three inputs) scores 3/3. Run the held-out suite, a different set of tests you wrote before the agent started and never showed it, and the same solution scores 0/3. If the agent can reach the stop condition, it’s a target, not a measurement.
  • The plausible screenshot: A grey, clickable button photographs beautifully. And Playwright’s toHaveScreenshot assertion goes green on its second run with no product change at all, because the first run wrote the baseline. A check that writes its own expected value can’t fail.

The common thread: distrust any evidence whose author and judge are the same process. In the second and third examples that’s literal. The agent can write what’s being checked, and the check writes its own expected value. In the first, the pipeline reports on itself, and its status stands in for the checker’s.

Protecting the oracle

An agent has two routes to a green suite. It can fix the code, or it can change the test. The second one has a sneaky cousin: change the configuration. That might be a jest.config exclusion that skips a failing file, a lowered coverage floor, or an eslint-disable comment that silences a rule.

Here’s how to close those routes:

  • Freeze the tests: Deny edits to test files and snapshot baselines with permission rules (settings that allow or deny specific tool calls), not just with an instruction. An instruction is a request. A permission rule is much stronger, because the harness enforces it whatever the model decides. But it isn’t airtight: a deny on Edit alone doesn’t stop a script or shell command from writing the same file. For anything that must hold, add a hook or the sandbox.
  • Watch the protected paths: In CI, flag any pull request that changes tests, baselines, or test configuration alongside the implementation.
  • Ratchet the suppressions: Count eslint-disable, @ts-expect-error, skipped tests, and assertions. The number of escape hatches shouldn’t go up. The number of assertions shouldn’t go down.
  • Gate on diff-scoped coverage: Coverage normally tells you how much of the whole codebase your tests execute. Diff-scoped coverage asks a narrower question: did any test run the lines this change touched? Across 4,882 agent-authored pull requests, 64.8% of the Python ones had no changed line executed by any existing test. In the Java pull requests that didn’t improve test coverage, agents deleted tests 2.6× as often as they added them. A green suite that never runs the changed lines is no evidence at all.
  • Test the judge: Feed your oracle deliberately broken code. Any broken variant (a mutant) that still passes is a blind spot. Do the same for a model-based grader: give it one wrong answer, and one right answer phrased three different ways.

Freezing is a permissions job, and some of the watching is a hook job. We’ll sort out which rule needs which tool in The Enforcement Ladder.

Match the evidence to the claim

Different claims need different proof:

ClaimWeak evidenceBetter evidence
It rendersThe agent says soA screenshot that a person or tool inspects
It looks rightA screenshotA diff against a baseline written by a different run
It worksA screenshotRole, name, and state assertions (the Save button exists and is disabled)
It’s doneTests passEach acceptance criterion mapped to evidence, re-run after the final edit
It’s deliveredIt mergedMerged, CI green on main after the merge, review threads resolved

A screenshot proves that a render happened. It doesn’t prove the change is correct, accessible, or responsive. Put role-and-state assertions (checks on what the page exposes, like a button’s accessible name and whether it’s disabled) in the agent’s inner loop, the edit-run-check cycle it repeats while it works. Save the pixel diffs for the merge boundary, the point where a change is about to land on main. And leave comparing two images to a pixel-diff tool or a person, not the model.

A claim without a check that could have failed isn’t a claim yet. It’s a hope.

Last modified on .