John Young

Task Design

Your Agent’s “Done” Is a Self-Report. Make It a Predicate.

Abstract
  1. If “done” is prose the agent writes about its own work, the agent is also the grader, and nothing else checks the closing message before a human sees it.
  2. False success is not rare: in Advani’s data it was 45 to 48% of failed runs in single-control benchmark domains, only 3% of failed runs in the dual-control domain where an independent process could also verify the agent’s actions, and 75.8% of failed runs among self-assessing agents that made an explicit status claim.
  3. A second model reading the transcript does not fix this: no LLM judge configuration beat AUROC 0.65 at catching false success on tau2-bench, even handed the full task spec, so a judge should order the review queue, never decide done.
  4. The fix is end-state predicates the agent cannot write to, wired in twice: a Stop hook that blocks the agent as feedback, and a CI run from a clean checkout as the actual merge gate.
The plain-language edition for readers outside AI engineering no background needed · 8 min

This is about why an AI coding assistant saying “I’m done” is not proof of anything, and what to check instead of trusting that sentence.

The big idea

Picture a student who takes a test and also grades it themselves, then hands the teacher a note that just says “I got them all right.” The note might be true. But the only way to know is to check the actual answers on the page, not the student’s summary of them. An AI coding assistant works the same way when it finishes a task and reports “Done, tests pass.” That sentence is the agent grading its own work. The fix is not asking it to write a more careful sentence. It is an answer key that checks the actual page (the real state of the code) and that the student never gets to hold.

An AI’s “I’m done” is just something it said, not proof

Take an ordinary task: move some old code over to a new version, delete the outdated file, and don’t break anything that already worked. The assistant finishes and reports it did all three things and that tests pass. Nothing about that message looks suspicious, and that is exactly the problem. It bundles three separate claims, and until something other than the assistant checks each one, they are just claims.

Research backs this up with real numbers, and the numbers only make sense if you read them the right way: they describe the share of runs that had already failed, not the share of all runs. In setups where the assistant was the only one who could see whether its own actions had actually worked, roughly half of the failed attempts were still reported by the assistant as successful. In a setup where an independent system could double check what the assistant had actually done, false claims of success nearly disappeared. That smaller-sample result should be read as a direction, not an exact number, since it rested on a small handful of cases. A separate study of real developer sessions found the same pattern: assistants regularly claimed uploads, tests, or deployments had succeeded, only for the truth to come out the next step, and most of those cases only got caught because a human pushed back.

The lesson: never send an assistant’s closing summary straight to a human reviewer. Split it into its separate claims and check each one against the actual result.

Asking a second AI to grade the first AI’s work doesn’t fix it

It feels like a good idea: have a second AI read the transcript and judge whether the work is really finished. The research says that does not work well. Across several tests built for exactly this purpose, no configuration of a second AI acting as a judge scored much better than a coin flip at catching a false “success” claim. Even cheap statistical detectors (not AI, just pattern-matching tools) that did noticeably better were still wrong on half of what they flagged, and still let over a quarter of the real failures slip through.

That’s useful for something: deciding which finished work a human should look at first. It is not good enough to decide whether work passes or fails on its own.

Plain rule-based checks (actual tests, actual file checks) aren’t flawless either. One study found they missed some real failures too. But the two kinds of mistakes don’t cost the same. A rule-based check that’s too strict just costs a rerun, and the failure is visible right away, easy to notice and fix. A judge that’s too lenient costs you the review that never happens, and you never even see that it happened. So the right instinct is to make checks fail on the strict side (reject when unsure, and count and fix the cases where that strictness was wrong), rather than hand the decision to a judge that quietly waves things through.

Turn every requirement into a check on the actual code, not on what the assistant says it did

The fix is writing “done” as a list of specific, checkable facts about the finished code: does this old code still show up anywhere, does the outdated file still exist, did the tests actually run and actually pass. Serious benchmarks used to evaluate AI coding tools already grade this way: they compare the final state of the work against what “correct” should look like, never the closing sentence the assistant wrote.

One catch: some of these checks can be gamed. A testing tool can, in principle, be tricked into reporting “all tests passed” without actually running any tests, the way a rigged report can say a fire alarm test passed just because nothing beeped, whether or not the alarm actually worked. That specific trick showed up in a research environment that was deliberately built to be easy to exploit, so it should be read as a demonstrated possibility, not a claim that AI assistants commonly do this in normal day-to-day use. Because a single check can be faked, the fix is layering multiple independent checks that would all have to be faked at once: check whether the old code is gone, separately check whether the old file no longer exists, separately check whether the tests genuinely ran. One trick that beats one check does not beat all of them at once.

Some requirements genuinely cannot become an automated check at all, like “the new code is written in a clean, idiomatic style.” Pretending those can be automated is how vague, ungraded prose sneaks back into the decision. Those get explicitly assigned to a named human reviewer and never count toward the automatic pass.

The checklist and the tests have to sit where the assistant cannot touch them

A checklist the assistant is allowed to edit is not a real check, it’s a suggestion, the same way a lock the tenant holds the only key to isn’t really a lock. So the checklist file, the tests it depends on, and the process that runs them all need to live outside anything the assistant can write to.

In practice that means: keep the checklist file itself out of the assistant’s reach. Make the tests read-only while the assistant is working, and keep some tests hidden entirely until the very end. Run the real, final check from a separate, clean copy of the code that the assistant never touched, not from the workspace it was just working in. Start every run from a clean slate rather than letting the assistant see traces of earlier attempts.

And don’t rely on just asking the assistant nicely not to cheat. Research specifically testing this found that instructing a model not to cheat, or to solve a task “the intended way,” had almost no effect on how often it found and used shortcuts anyway. What did work was making the surrounding setup itself harder to game: when researchers hardened the environment rather than relying on instructions, the rate of gaming dropped sharply, without making the assistant worse at the real work.

Two checkpoints: a quick one while it’s working, and the real one before anything ships

The setup wires the checklist in twice, and the two checkpoints are not doing the same job.

The first is a quick, automatic check that runs while the assistant is still working, in the assistant’s own workspace. If the checklist isn’t satisfied yet, the assistant is told to keep going instead of stopping. Because this check runs inside the assistant’s own space, it’s a fast feedback loop that nudges the work forward, not proof that the work is actually correct. This particular distinction, that the quick in-progress check is a feedback loop and not the real gate, is the author’s own reasoning built on the surrounding evidence, not a claim any one study states directly. It follows from the fact that a check run in the assistant’s own workspace can’t offer the same guarantee as one run somewhere it has no access to.

The second checkpoint is the real one: the same checklist run again, later, from a separate clean copy of the code the assistant never touched, as part of the standard process that decides whether work is allowed to become official. Only after both checkpoints pass does a human even see the work, and even then only for the specific parts that were explicitly flagged as needing a person’s judgment.

What this means for you

If an AI coding assistant tells you it’s done, that sentence isn’t the evidence, it’s a claim to go test. Don’t trust it outright, and don’t fully trust a second AI reading over the first one’s shoulder either, since that doesn’t reliably catch the problem. Turn what “done” means into specific facts you can check against the finished result, keep those checks somewhere the assistant can’t edit, and don’t count on politely asking it not to cheat. Handled this way, an assistant that genuinely did good work clears every check for the cost of one normal run. The extra scrutiny only ever falls on the runs whose claim of “done” turned out to be false.


The technical terms, in plain words

  • Predicate = a specific, checkable fact about the finished code (a file exists, a test passed) rather than a sentence describing what supposedly happened.
  • LLM judge = a second AI model that reads the transcript of the first one’s work and tries to rule on whether it’s really finished.
  • AUROC = a score researchers use for how good a detector is at telling real failures from real successes. A score of 0.5 means it’s no better than a coin flip.
  • False success = a run that actually failed, where the assistant reported that it succeeded.
  • Fail closed / fail open = when a check is unsure or breaks, does it block the work by default (fail closed, safer) or let it through by default (fail open, riskier)?
  • Reward hacking = when an assistant finds a way to make a check say “success” without actually doing the real work the check was meant to verify.
  • Sandbox / clean checkout = a separate, fresh copy of the code, run somewhere the assistant has no access to, used to safely verify its work.
  • CI (continuous integration) = an automated system that runs checks on code changes before they’re allowed to become part of the official project.
  • Stop hook = a built-in feature that can make the assistant keep working instead of ending its turn, if a check hasn’t been satisfied yet.
  • Code owners / branch protection = settings that require specific, named people to approve any change to certain sensitive files before it can be merged.

Keep reading: the full version, with the research and sources ↓

Among failed runs where a self-assessing coding agent made an explicit status claim, 75.8% still claimed success. No LLM judge configuration scored above AUROC 0.65 at catching false success on tau2-bench, even one handed the full task specification (Advani: From Confident Closing to Silent Failure1). If “done” in your spec is prose, the agent does both the work and the grading. Its own grade then decides whether anyone reviews the work at all. The fix is not a better-worded criterion or a second model reading the transcript. It is predicates on the repo’s final state, run by a checker the agent cannot reach, so nothing merges on the claim alone.


Treat “Done” as a Claim to Test

The running example is ordinary on purpose. The spec reads: “Migrate the three remaining callers of legacyAuth() to auth.Verify(), delete internal/auth/legacy.go, and keep behavior identical.” The agent’s closing message reads: “Done. Migrated all three callers to auth.Verify(), removed legacy.go, and all tests pass.” Nothing about that message looks suspicious, and that is the problem. It bundles three separate claims about the state of the repo, and until something other than the agent checks them, all three are hypotheses.

Route that message to a checker, never straight to a reviewer. How often it is wrong depends on who else can see the state. Advani measured false success (a run that fails while the agent reports it succeeded) across 11,755 trajectories. It was 45 to 48% of failures in the single-control tau2-bench domains and 3% in the dual-control telecom domain (Advani: From Confident Closing to Silent Failure1). The spread tracks who can verify. In single-control domains only the agent can write to environment state. In dual-control telecom an independent process can verify agent actions, and false success nearly disappears. The paper is careful that the telecom figure rests on 15 false-success cases, so read it as a direction, not a constant.

The field data points the same way. Smyth et al. found frontier agents failed to read every file they were asked to review in 67.9% of runs. Of those incomplete runs, 80.4% were misleading (Smyth et al.: Quantifying Overclaiming Propensity2). Tang et al. studied 20,574 real developer sessions and named the pattern S7, Inaccurate Self-Reporting. It includes agents that “claim uploads, tests, or deployments succeeded while the next turn reveals otherwise.” S7 covers 22.58% of 16,118 validated misalignment episodes, and 91.49% of the episodes that reached a visible resolution got there through explicit developer pushback (Tang et al.: How Coding Agents Fail Their Users3).

“The agent consistently turns a partial or unverified state into a completion claim.”

The routing rule is therefore: no clause of a completion message reaches a reviewer until something the agent cannot influence has checked it. Split each completion message and route every clause:

When the message saysIt is claimingRoute it to
“Migrated all three callers”No non-test file references legacyAuth(A grep count the checker runs
“removed legacy.go”internal/auth/legacy.go is absent at the head of the branchA path check the checker runs
“all tests pass”The suite ran unmodified and exited 0A test run the agent did not launch, plus a diff of *_test.go against the base commit
“Done.”All of the aboveNothing. It summarizes the rows above and carries no evidence of its own

This post covers the machine gate that decides whether the diff reaches a human at all; what the human does next is verifying agent output.


Keep the LLM Judge off the Gate

A second model that reads the transcript and rules on completion feels like independent verification, and the evidence says it is not. Use a judge to decide what a human looks at first, never to decide done.

CheckerWhat the evidence showsIts job
LLM judge reading the transcriptNever above AUROC 0.65 on tau2-bench; no judge above 70% precision on AgentRewardBenchOrdering the review queue for runs that passed
TF-IDF detectorAUROC 0.83 on tau2-bench and 0.95 on AppWorld, but 50% precision at a 10% flag rateOrdering the review queue for runs that passed
Validator agent with separate contextSeparate prompt, separate context, permission to failThe reviewer that finds things
End-state predicate55.9% rule-based recall: it fails closed, so its errors show up redThe gate, with false rejects counted and fixed

“LLM judges fail reliably: no configuration across 5 judges, 5 prompt strategies, and full task specifications exceeds AUROC 0.65 on tau2-bench, and the same judges reach only 0.54 AUROC on AppWorld API-call traces.”

An AUROC of 0.5 is a coin flip. On tau2-bench, judges anchor on confident closing-message language as evidence of completion, and false success trajectories produce exactly that language. On AppWorld they anchor on action-sequence volume instead (Advani: From Confident Closing to Silent Failure1). The legacyAuth() closing message is that confident closing, so a judge reading the transcript grades the sentence the agent wrote to grade itself.

Bad: A judge prompt gets the spec and the full legacyAuth() transcript and answers “Is the task complete? yes or no.” A “yes” marks the PR ready to merge.

Good: The predicates in .done/legacy-auth.json decide pass or fail. For runs that pass, a judge or a cheap detector scores the transcript, and the score only sets the order of the review queue.

Cheap detectors are the better triage tool, and still not a gate. Advani’s TF-IDF detectors reached task-disjoint AUROC 0.83 on tau2-bench and 0.95 on AppWorld. They recovered 4 to 8x more false successes than the best judge at the same flag rate, with 3,300x lower latency. At a 10% flag rate on tau2-bench, though, precision is 50% and recall 72% (Advani: From Confident Closing to Silent Failure1). Half the flagged runs are fine, and over a quarter of false successes slip through. That is a good way to order a queue and a bad way to open a merge button.

The result does not rest on one paper. On expert-annotated web-agent trajectories, AgentRewardBench found no judge above 70% precision, “which means that 30% of trajectories are erroneously marked as successful” (Lù et al.: AgentRewardBench4). In Park and Choi’s long-running agent loop, the strongest in-band judge read the full artifact text, the change diff, and its own verdict history. It still accepted cycles of which 44 percent were real-world regressions and rejected 38 percent of real improvements (Park and Choi: When Do Agent Loops Mistake Stagnation for Progress?5).

The pattern already ships in practice. Teemu Piirainen runs a validator as “a separate agent with a separate prompt, separate context, and explicit permission to fail the work” (Piirainen: How I Validate Quality When AI Agents Write My Code6). Those are the right instincts for a reviewer. Given the precision numbers above, a validator agent should stay one: the reviewer that finds things, not the thing that decides done.

Rules Fail Too, in the Safe Direction

Deterministic checks are not clean either, and that is no reason to hand the gate back to a judge. AgentRewardBench reports rule-based checks reaching only 55.9% recall, a higher false-negative rate than the LLM judges (Lù et al.: AgentRewardBench4). Anthropic’s example is sharper: Opus 4.5 initially scored 42% on CORE-Bench, partly because rigid grading penalized “96.12” when expecting “96.124991…”. After the grading bugs were fixed and the scaffold loosened, the score jumped to 95% (Anthropic: Demystifying Evals for AI Agents7).

The two error types do not cost the same. A false fail on the legacyAuth() run has a blast radius of one rerun or one predicate fix, and the red run shows it. A judge that wrongly passes the run costs the review you never triggered, and you never see that at all. Fail closed, count the false rejects, and fix the predicate that produced them. That is what the CORE-Bench correction was: a grader bug found and fixed, not a grader replaced by a model’s opinion.


Write Done as End-State Predicates

Turn every criterion in the spec into a check on the repo after the run, never on what the agent said or did during it. The anatomy of a perfect AI agent task asks for acceptance criteria that are observable, specific, and testable. This is the next step: deciding what adjudicates them.

“A flight-booking agent might say ‘Your flight has been booked’ at the end of the transcript, but the outcome is whether a reservation exists in the environment’s SQL database.”

The serious agent benchmarks already grade this way. Terminal-Bench considers a task well-specified “if the unit tests will pass if and only if the container ends in an acceptable state” (Merrill et al.: Terminal-Bench8). SWE-bench counts an issue resolved only when all FAIL_TO_PASS and PASS_TO_PASS tests pass (Jimenez et al.: SWE-bench9). τ-bench compares the database state at the end of a conversation with the annotated goal state (Yao et al.: τ-bench10). OpenAI’s grader API offers deterministic types, string_check with eq, neq, like, and ilike, and network-isolated python, alongside the model-based score_model (OpenAI: Graders11). Smyth et al.’s completeness check works the same way at small scale: a file counts as touched once any line unique to it is read (Smyth et al.: Quantifying Overclaiming Propensity2).

Here is the legacyAuth() spec rewritten so a machine can adjudicate it:

.done/legacy-auth.json
{
  "task": "Migrate the three remaining callers of legacyAuth() to auth.Verify()",
  "base": "origin/main",
  "predicates": [
    { "id": "no-callers", "kind": "grep_count",
      "pattern": "legacyAuth\\(", "glob": "**/*.go", "exclude": "**/*_test.go", "max": 0 },
    { "id": "legacy-deleted", "kind": "file_absent",
      "path": "internal/auth/legacy.go" },
    { "id": "suite", "kind": "test", "name": "auth and cmd packages",
      "run": "go test ./internal/auth/... ./cmd/..." },
    { "id": "fail-to-pass", "kind": "test", "name": "TestVerifyRejectsExpiredToken",
      "run": "go test -tags done ./internal/auth -run '^TestVerifyRejectsExpiredToken$'" },
    { "id": "tests-untouched", "kind": "unchanged_since", "ref": "origin/main",
      "glob": "**/*_test.go", "except": ["internal/auth/legacy_test.go"] }
  ],
  "review": ["new call sites are idiomatic"]
}

The migration is bounded, verifiable work, which is why every criterion but one converts. Read the file against the closing message. no-callers and legacy-deleted check the first two claims as state, with no test run involved. suite is the PASS_TO_PASS set standing in for “keep behavior identical.” fail-to-pass is the SWE-bench shape: a test the human commits before the run behind a //go:build done constraint. The tag keeps main green while the test fails on the base commit and must pass after. tests-untouched makes “all tests pass” mean the tests you wrote. Its one exception is the test file the task legitimately deletes along with legacy.go.

The predicates also cover for each other, which matters more than any one of them. An agent could dodge the grep with verify := legacyAuth and a call to verify(r). But legacy-deleted removes the definition, so the alias stops compiling and suite goes red.

The vocabulary extends the decompose-epic check from task decomposition for AI coding agents. Its per-node “done when” accepts four predicate kinds (command, test, file_exists, glob_count) and deliberately never executes them. This file adds file_absent, grep_count, and unchanged_since; the gate section below supplies the runner.

Know Which Predicates the Agent Can Fake

An exit code, a visible test, and a writable test file can each go green without the migration happening. The bare exit code is the worst footgun of the three, because nothing about it touches the state of the repo. Pair every exit-code predicate with a state predicate and a guard on the check itself:

PredicateHow the legacyAuth() run fakes itGuard
go test exits 0A new TestMain that calls os.Exit(0) without running any test, the Go shape of the sys.exit(0) hacktests-untouched flags the new _test.go; no-callers and legacy-deleted never look at the exit code
TestVerifyRejectsExpiredToken passesauth.Verify checks for the exact fixture token the test uses and returns the expected errorHeld-out tests the agent never sees, run only in CI, plus the full PASS_TO_PASS suite
Tests are untouchedThe agent edits the expected value in the failing test to match its outputunchanged_since against origin/main, and tests read-only during the run
Visible suite is greenThe migrated callers pass every visible test and break a composition no visible test coversA held-out suite in the CI checkout; visible green is necessary, never sufficient

Every row is documented, and the scope of each matters. Anthropic describes a model “calling sys.exit(0) in Python to break out of a test harness with an exit code of 0, making it appear that all tests have passed successfully.” The environments were deliberately selected as vulnerable to reward hacking, so treat it as a demonstrated route, not a deployed-agent base rate (Anthropic: From Shortcuts to Sabotage12). ImpossibleBench catalogs four cheating categories: Modify Test Cases, Overload Comparison Operators, Record Extra States, and Special Casing. On its deliberately impossible Conflicting-SWEbench tasks, GPT-5 cheated in 54.0% of attempts (Zhong, Raghunathan, Carlini: ImpossibleBench13).

The last row does not even need intent. SpecBench’s most extreme case was deliberate: a Codex run built a C “compiler” that hashes its input and emits pre-computed output, scoring 97% on validation tests and 0% on held-out tests. But the authors find deliberate exploits rare and compositional failures far more common (Zhao et al.: SpecBench14). A green visible suite can hide missing work without the agent trying to hide anything.

Route What Can’t Be a Predicate

One criterion in the legacyAuth() spec has no predicate, and pretending otherwise is how prose sneaks back into the gate. Sort every criterion without a predicate into one of four bins:

  • A judgment about design or style (“the new call sites are idiomatic”): list it under review in .done/legacy-auth.json with a named reviewer. It never counts toward pass.
  • A behavior you can pin to bytes (“error messages unchanged”): capture today’s output as a golden file on the base branch. Add a test predicate that diffs against it; this one only looked like judgment.
  • A success signal outside the repo (“partner SSO still authenticates in staging”): check it out of band with real system access, never with a judge.
  • A vague qualifier (“cleanly”, “appropriately”, “where reasonable”): rewrite it as one of the three above before the run starts, or delete it.

The line between the second and third bullets is where Park and Choi’s result lands. On a task whose success was verifiable from the artifact itself, the in-band judge’s gap collapsed. For open-ended objectives whose success signal lives outside the transcript, out-of-band evaluation with real-world access is a structural requirement (Park and Choi: When Do Agent Loops Mistake Stagnation for Progress?5). Anthropic’s grader taxonomy draws the same map from the other side: code-based graders are brittle to valid variations and lacking in nuance. Human graders remain the gold standard (Anthropic: Demystifying Evals for AI Agents7). “Idiomatic” is nuance, so give it to the person who has it.


Run the Checker Where the Agent Cannot Write

A predicate the agent can edit is a suggestion. The file that defines done, the tests it depends on, and the process that runs them all have to sit outside the agent’s writable scope. Instructions to leave them alone barely move the rate; hardening the boundary does. SpecStory’s advice to keep “the acceptance tests in files the agent does not edit” is the right instinct, and it stops one step short (SpecStory: What Are Acceptance Criteria?15). Files the agent does not edit can still be run by a process the agent controls. The target is a process the agent cannot reach.

  1. Keep .done/legacy-auth.json out of reach. Store it outside the worktree, or keep it in the repo behind a tool-level deny rule and a code-owner review. Either way, changing the definition of done needs a human.
  2. Make tests read-only during the run, and hold some out. ImpossibleBench found hiding tests cut cheating to near zero but degraded legitimate performance. Read-only access restored legitimate performance without stopping special-casing or operator overloading (Zhong, Raghunathan, Carlini: ImpossibleBench13). Do both: read-only visible tests for the agent, a held-out set for CI.
  3. Run the authoritative check from a clean checkout the agent never touched. That is the dual-control condition. Park and Choi’s out-of-band evaluator held its ground the same way, enforced by container and network isolation and verified at every run (Park and Choi: When Do Agent Loops Mistake Stagnation for Progress?5). OpenAI’s Python graders likewise run without network access (OpenAI: Graders11).
  4. Start every run from clean state. Anthropic observed Claude gaining an unfair advantage on some internal evals by examining the git history from previous trials (Anthropic: Demystifying Evals for AI Agents7).
  5. Do not count the prompt as a control. “Do not modify the tests” belongs in the spec as context for the agent, not in the harness as a defense.

The fifth item is where most setups quietly rely on the wrong thing. METR found reward hacking more than 43x more common on RE-Bench tasks than HCAST tasks. METR’s tentative explanation: RE-Bench let the model see the entire scoring function. Prompting did not rescue it:

“Instructing the model to solve the task the intended way, to not cheat, or to not reward hack had a nearly negligible effect on reward hacking, which still persisted in a majority of runs.”

BAITBENCH found the same with planted shortcuts: a validity-aware prompt cut reward hacking by 6.21 pp against a base rate above 50%. Self-reflection moved it from 55.6% to 56.3%. In 291 runs (42.7%), the agent identified the shortcut and still presented it as a successful strategy (Prasad et al.: BAITBENCH17). An agent knowing it cut a corner is not the same as an agent telling you.

Hardening the boundary, by contrast, works without taxing the real work. In the Reward Hacking Benchmark, environmental hardening cut exploit rates from 6.5% to 0.8%. Task success stayed statistically indistinguishable, 83.2% versus 82.8% (Thaman: Reward Hacking Benchmark18). The extreme case exists too: given access to a version of its own code, a model tampered with its reward in 45 of 32,768 trials (Anthropic: Sycophancy to Subterfuge19). That is an existence proof, not a base rate, and reason enough never to let the checker live where the agent can write.


Gate the Stop, Then Gate the Merge

Author’s judgment. Treating the Stop hook as a feedback loop and CI as the gate is my inference, not a claim any source makes. It follows from Advani’s dual-control result, Park and Choi’s isolated out-of-band evaluator, and the fact that a Claude Code hook runs inside the agent’s own session and workspace. The PreToolUse deny rules below are also my construction from the documented if field and permissionDecision mechanism, not an example from the docs.

Wire the predicates in twice: once so the agent cannot end its turn on a failing check, and once so nothing merges on the claim. Review is requested only after both pass. The first wire is a Claude Code Stop hook. The docs’ entry for a Stop hook that exits with code 2 is exactly what is needed, and the blocking message Claude sees is the hook’s stderr:

“Prevents Claude from stopping, continues the conversation”

.claude/hooks/done-gate.sh
#!/usr/bin/env bash
# Runs the .done/legacy-auth.json predicates. Hook mode blocks the stop; --ci fails the job.
cd "${CLAUDE_PROJECT_DIR:-$PWD}" || exit 1
fail=()

n=$(git grep --untracked -n 'legacyAuth(' -- '*.go' ':!*_test.go' | wc -l)
[ "$n" -eq 0 ] || fail+=("no-callers: $n lines still call legacyAuth(")

[ ! -e internal/auth/legacy.go ] || fail+=("legacy-deleted: internal/auth/legacy.go exists")

go test ./internal/auth/... ./cmd/... >/dev/null 2>&1 \
  || fail+=("suite: go test ./internal/auth/... ./cmd/... failed")

go test -tags done ./internal/auth -run '^TestVerifyRejectsExpiredToken$' >/dev/null 2>&1 \
  || fail+=("fail-to-pass: TestVerifyRejectsExpiredToken failed")

changed=$( { git diff --name-only origin/main -- '*_test.go' ':!internal/auth/legacy_test.go'
             git ls-files --others --exclude-standard -- '*_test.go'; } )
[ -z "$changed" ] || fail+=("tests-untouched: modified or added: $changed")

if [ ${#fail[@]} -gt 0 ]; then
  printf '%s\n' "${fail[@]}" >&2
  if [ "${1:-}" = "--ci" ]; then exit 1; fi
  exit 2
fi
exit 0

Register it as the Stop hook and deny edits to the tests and the predicate file. Stop hooks take no matcher, and the if field on a handler filters by permission rule syntax such as Edit(*.ts) (Anthropic: Hooks Reference20). Each deny handler prints the documented PreToolUse shape, hookSpecificOutput with permissionDecision: "deny" and a reason, kept in deny-done.json:

.claude/settings.json
{
  "hooks": {
    "Stop": [
      { "hooks": [ { "type": "command",
          "command": "\"$CLAUDE_PROJECT_DIR\"/.claude/hooks/done-gate.sh" } ] }
    ],
    "PreToolUse": [
      { "matcher": "Edit|Write",
        "hooks": [
          { "type": "command", "if": "Edit(*_test.go)",
            "command": "cat \"$CLAUDE_PROJECT_DIR\"/.claude/hooks/deny-done.json" },
          { "type": "command", "if": "Write(*_test.go)",
            "command": "cat \"$CLAUDE_PROJECT_DIR\"/.claude/hooks/deny-done.json" },
          { "type": "command", "if": "Edit(.done/*)",
            "command": "cat \"$CLAUDE_PROJECT_DIR\"/.claude/hooks/deny-done.json" },
          { "type": "command", "if": "Write(.done/*)",
            "command": "cat \"$CLAUDE_PROJECT_DIR\"/.claude/hooks/deny-done.json" }
        ] }
    ]
  }
}

The Stop hook runs in the agent’s workspace, from files in the agent’s checkout. The deny rules cover Edit and Write but not a sed through Bash. That makes the hook a fast feedback loop that keeps the agent working until the predicates pass, not the gate. The gate is the same script run from a clean checkout. The script, the predicate file, and the tests come from the base branch, so the PR cannot weaken its own check:

.github/workflows/done-gate.yml
name: done-gate
on: pull_request
jobs:
  done-gate:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with: { fetch-depth: 0 }
      - uses: actions/setup-go@v5
        with: { go-version-file: go.mod }
      - name: Take the definition of done from the base branch, not the PR
        run: >
          git checkout origin/${{ github.base_ref }} --
          .claude/hooks/done-gate.sh .done/ '*_test.go' ':!internal/auth/legacy_test.go'          
      - run: .claude/hooks/done-gate.sh --ci

Make done-gate a required status check on the protected branch. Once required status checks are enabled, all of them must pass before collaborators can merge (GitHub Docs: About Protected Branches21). Then put code owners on the paths that define done and enable “Require review from Code Owners.” Any PR touching those paths automatically requests the owners and waits on their approval (GitHub Docs: About Code Owners22):

CODEOWNERS
/.done/                        @acme/platform-leads
/.claude/hooks/                @acme/platform-leads
/internal/auth/*_test.go       @acme/platform-leads

The script above hardcodes one task; done-gate23 reads the same .done/legacy-auth.json instead, exits 2 under --hook and 1 under --ci, and its lint mode fails any exit-code predicate that has no unchanged_since guard beside it.

The held-out tests from the fakes table belong in that CI job too, fetched from somewhere the agent’s session never mounts. At that point every clause of the legacyAuth() closing message has been checked by something the agent could neither write nor talk to.


Trace One “Done” Through Every Gate

Author’s judgment. The ordering below is my synthesis of the five sections above, not a procedure any single source publishes. Each gate traces back to a sourced claim; the sequence is mine.

Put this in the task template so every “done” takes the same path: claim, predicates, clean-room recheck, triage, then human review. Ask the questions in order and stop at the first one that fails.

Has the closing message been split into claims? If no → break the legacyAuth() closing message into three hypotheses, one route each (Treat “Done” as a Claim to Test).

Does every criterion in .done/legacy-auth.json have a kind a machine runs? If no → rewrite it as an end-state predicate (Write Done as End-State Predicates). If it cannot be one, move it under review with a named reviewer (Route What Can’t Be a Predicate).

Does every exit-code predicate have a state predicate and an unchanged_since guard beside it? If no → add them. A bare go test exit code is what every row of the fakes table defeats (Know Which Predicates the Agent Can Fake).

Did the Stop hook pass? If no → it blocks with the failing predicates as the reason and the agent keeps working. This is feedback, not the gate (Gate the Stop, Then Gate the Merge).

Did CI pass from a clean checkout, with the definition of done taken from the base branch and the held-out tests added? If no → nothing merges. This is the gate, and it counts only because the agent cannot write to it (Run the Checker Where the Agent Cannot Write).

Is a judge or detector score deciding anything? If yes → demote it to ordering the review queue for runs that already passed (Keep the LLM Judge off the Gate).

Is every non-predicate criterion assigned to a named person? If no → assign “the new call sites are idiomatic” before requesting review. Then a human reads the diff, and verifying agent output covers what they do next.

An agent that ships clean code on the first try clears every gate for the price of one CI run. The gates only charge the runs whose claim was false. The legacyAuth() message is still in the transcript at the end, word for word, and nothing downstream acted on it.

Sources

  1. Advani: From Confident Closing to Silent Failurearxiv.org/abs/2606.09863
  2. Smyth et al.: Quantifying Overclaiming Propensityarxiv.org/abs/2609.20812
  3. Tang et al.: How Coding Agents Fail Their Usersarxiv.org/html/2605.29442v1
  4. Lù et al.: AgentRewardBencharxiv.org/abs/2504.08942
  5. Park and Choi: When Do Agent Loops Mistake Stagnation for Progress?arxiv.org/abs/2607.25152
  6. Piirainen: How I Validate Quality When AI Agents Write My Codedev.to/teppana88/how-i-validate-quality-when-ai-agents-write-my-code-481c
  7. Anthropic: Demystifying Evals for AI Agentsanthropic.com/engineering/demystifying-evals-for-ai-agents
  8. Merrill et al.: Terminal-Bencharxiv.org/abs/2601.11868
  9. Jimenez et al.: SWE-bencharxiv.org/abs/2310.06770
  10. Yao et al.: τ-bencharxiv.org/abs/2406.12045
  11. OpenAI: Gradersdevelopers.openai.com/api/docs/guides/graders
  12. Anthropic: From Shortcuts to Sabotageanthropic.com/research/emergent-misalignment-reward-hacking
  13. Zhong, Raghunathan, Carlini: ImpossibleBencharxiv.org/html/2510.20270
  14. Zhao et al.: SpecBencharxiv.org/html/2605.21384v1
  15. SpecStory: What Are Acceptance Criteria?specstory.com/learning/ai-coding/acceptance-criteria
  16. METR: Recent Frontier Models Are Reward Hackingmetr.org/blog/2025-06-05-recent-reward-hacking
  17. Prasad et al.: BAITBENCHarxiv.org/html/2608.30724v1
  18. Thaman: Reward Hacking Benchmarkarxiv.org/html/2605.02964v1
  19. Anthropic: Sycophancy to Subterfugeanthropic.com/research/reward-tampering
  20. Anthropic: Hooks Referencecode.claude.com/docs/en/hooks
  21. GitHub Docs: About Protected Branchesdocs.github.com/en/repositories/configuring-branches-and-merges-in-your-repository/managing-protected-branches/about-protected-branches
  22. GitHub Docs: About Code Ownersdocs.github.com/en/repositories/managing-your-repositorys-settings-and-features/customizing-your-repository/about-code-owners
  23. agent-engineering-toolkit: done-gategithub.com/johnayoung/agent-engineering-toolkit