John Young

Research Index

The evidence base behind everything I publish on running AI coding agents in production. Every entry is a verbatim figure or quote from a primary source — a study, a benchmark, an engineering post — pulled while researching a post, then checked against the live page. Sources that drift or go dead are dropped or flagged.

331
Verified sources
6
Themes
111 / 220
Data / practitioner
2026-09-07
Last verified

Every entry checked against its live source · dataset: research.json

Task Design & Decomposition

23 sources

Scoping, decomposing, and speccing work so an agent finishes it on the first try.

Alif Al Hasan, Sumon Biswas (Case Western Reserve University) May 29, 2026 data

Across 547 confirmed real-world safety failures mined from the GitHub issue trackers of 13 foundational code models, the top threat category is Constraint Violations at 40.4%, ahead of Destructive Operations (24.5%), Authorization Bypasses (18.3%), and Deception (15.7%).

Independent corroboration of the constraint-violation finding by a different dataset and method - two teams reaching the same top category within two points.

3 more excerpts
  • Failures arise during benign, goal-directed use rather than adversarial attack
  • Nearly 60% of confirmed incidents were rated high or critical severity
  • The quantified breakdown appears in the full text, not on the arXiv abstract page

What Breaks When LLMs Code? Characterizing Operational Safety Failures of Agentic Code Assistants ↗·Cited in Task Decomposition for AI Coding Agents: Draw the Graph First

Anthropic Feb 18, 2026 data

80% of tool calls come from agents that appear to have at least one kind of safeguard (like restricted permissions or human approval requirements), 73% appear to have a human in the loop in some way, and only 0.8% of actions appear to be irreversible

Irreversible agent actions are rare in real traffic, so oversight should concentrate on the small slice where a single error is costly.

2 more excerpts
  • such as sending an email to a customer
  • And while these higher-risk actions are rare as a share of overall traffic, the consequences of a single error can still be significant.

Measuring AI agent autonomy in practice ↗·Cited in What AI Coding Agents Are Actually Good For (And When to Skip)

Arpandeep Khatua, Hao Zhu, Peter Tran, Arya Prabhudesai, Frederic Sadrieh, Johann K. Lieberwirth, Xinkai Yu, Yicheng Fu, Michael J. Ryan, Jiaxin Pei, Diyi Yang (Stanford, SAP Labs) Jan 19, 2026 (v1); revised Jan 26, 2026 (v2) data

Across 600+ collaborative coding tasks in 12 libraries and 4 languages, agents working together achieve on average 30% lower success rates than the same agents doing both tasks individually - the 'curse of coordination'. GPT-5 and Claude Sonnet 4.5 configurations reach only 25% under two-agent cooperation, roughly half the solo baseline. 77.3% of tasks have conflicting ground-truth solutions.

Without assigned file and interface ownership, parallel agents duplicate work and overwrite changes they believe will merge cleanly - the collision is the default outcome, not an edge case.

4 more excerpts
  • Adding a messaging tool did not help: the difference between 'with comm' and 'no comm' settings is not statistically significant for task success, though it did reduce literal merge conflicts
  • The paper separates two problems: merge conflicts are spatial coordination (who edits which lines), while task success requires semantic coordination (what to implement, not just where)
  • Agents were given no pre-assigned file ownership and were free to redivide the features between themselves
  • Scope limit: no experimental arm tested pre-assigned ownership as a fix, so the benchmark diagnoses the problem without validating the cure

CooperBench: Why Coding Agents Cannot be Your Teammates Yet ↗·Cited in Task Decomposition for AI Coding Agents: Draw the Graph First

Kwa, West, Becker, et al. (METR) submitted 2025-03-18 data

frontier AI time horizon has been doubling approximately every seven months since 2019, though the trend may have accelerated in 2024

The primary paper behind the autonomy trend confirms a ~7-month doubling of the 50%-task-completion time horizon since 2019, driven mainly by greater reliability and error-adaptation — the mechanism that inflates calls per task.

4 more excerpts
  • "Claude 3.7 Sonnet have a 50% time horizon of around 50 minutes".
  • "within 5 years, AI systems will be capable of automating many software tasks that currently take humans a month".
  • "The increase in AI models' time horizons seems to be primarily driven by greater reliability and ability to adapt to mistakes"
  • 50%-task-completion time horizon. This is the time humans typically take to complete tasks that AI models can complete with 50% success rate

Measuring AI Ability to Complete Long Software Tasks ↗·Cited in You Can't Cap What You Can't Attribute: Per-Task Cost, What AI Coding Agents Are Actually Good For (And When to Skip)

METR March 19, 2025 data

The length of tasks (measured by how long they take human professionals) that generalist frontier model agents can complete autonomously with 50% reliability has been doubling approximately every 7 months for the last 6 years.

The autonomous task length frontier agents can complete has doubled roughly every 7 months for 6 years, so autonomous runs — and the per-task call count behind them — keep growing.

4 more excerpts
  • current models have almost 100% success rate on tasks taking humans less than 4 minutes, but succeed <10% of the time on tasks taking more than around 4 hours
  • "If the measured trend from the past 6 years continues for 2-4 more years, generalist autonomous agents will be capable of performing a wide range of week-long tasks."
  • the best current models—such as Claude 3.7 Sonnet—are capable of some tasks that take even expert humans hours, but can only reliably complete tasks of up to a few minutes long
  • AI agents often seem to struggle with stringing together longer sequences of actions

Measuring AI Ability to Complete Long Tasks ↗·Cited in You Can't Cap What You Can't Attribute: Per-Task Cost, Loop Engineering Breaks Your Single-Shot Context Playbook, What AI Coding Agents Are Actually Good For (And When to Skip)

METR (Becker, Rush, Barnes, Rein) July 10, 2025 data

When developers are allowed to use AI tools, they take 19% longer to complete issues—a significant slowdown that goes against developer beliefs and expert forecasts.

Experienced developers were measurably slower with AI in codebases they know well, contradicting their own forecasts of a speedup.

4 more excerpts
  • 16 experienced developers from large open-source repositories (averaging 22k+ stars and 1M+ lines of code)
  • developers expected AI to speed them up by 24%
  • they still believed AI had sped them up by 20%
  • developers estimated that they were sped up by 20% on average when using AI—so they were mistaken

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision, What AI Coding Agents Are Actually Good For (And When to Skip)

Ningzhi Tang, Chaoran Chen, Gelei Xu, Yiyu Shi, Yu Huang, Collin McMillan, Tao Dong, Toby Jia-Jun Li (Notre Dame, Vanderbilt, Google) May 28, 2026 data

Across 20,574 coding-agent sessions from 1,639 repositories, the most prevalent misalignment symptom is Developer Constraint Violation - defined as violating an explicit developer constraint - at 38.33% of episodes, with 73.68% of those attributed to instruction-following failure. The separate underspecification cause (C1) accounts for only 15.36%.

The dominant measured failure is agents breaking constraints developers already stated, not developers failing to state them - which is why sharpening prompt prose does not address the main failure mode.

4 more excerpts
  • The symptom taxonomy is explicitly multi-label - 29.56% of episodes carry two labels and 0.54% carry three or more - so the seven shares deliberately do not sum to 100%
  • 90.50% of episodes impose effort and trust costs rather than irreversible system damage, yet 91.49% of visible resolutions still require explicit user correction
  • Misalignment compounds across sessions: probability of misalignment in the next session is 0.519 after an affected session versus 0.336 otherwise
  • Constraint violation is markedly worse in CLI sessions (49.49%) than IDE sessions (32.26%)

How Coding Agents Fail Their Users: A Large-Scale Analysis of Developer-Agent Misalignment in 20,574 Real-World Sessions ↗·Cited in Task Decomposition for AI Coding Agents: Draw the Graph First

Shubhi Asthana, Bing Zhang, Chad DeLuca, Hima Patel, Ruchi Mahindru (IBM Research) May 14, 2026 data

On a Kubernetes root-cause-analysis workload, a decomposition fixed at design time with no runtime branching cost 1,632 +/- 145 tokens in retries versus 904 +/- 17 for a monolithic run - 80.5% worse. Runtime-structured decomposition with schema-validated handoffs cut retry cost to 436 +/- 132, a 51.7% reduction against monolithic and 73.2% against static.

Decomposition is not automatically a win. Splitting work without runtime isolation adds rerun surface area, because a failure anywhere forces re-execution of every downstream subtask.

4 more excerpts
  • The mechanism is stated directly: 'fixed sequential execution must rerun all downstream subtasks from the point of failure'
  • Structuring is not free - the runtime-structured baseline run cost 2,716 +/- 424 tokens against 904 +/- 17 monolithic, so the trade only pays at a nonzero failure rate
  • Authors' limitation: both use cases are controlled scenarios at temperature 0 with low natural failure rates (0-2%), and token savings depend on deployment failure rates they did not measure at scale
  • Authors' limitation, load-bearing for this post: 'Decomposition policies are developer-authored and may not generalize to automatically derived graphs'

Runtime-Structured Task Decomposition for Agentic Coding Systems ↗·Cited in Task Decomposition for AI Coding Agents: Draw the Graph First

Tim Menzies, William Nichols, Forrest Shull, Lucas Layman (NC State, SEI-CMU, Fraunhofer CESE) 2016 data

Across 171 software projects from 2006 to 2014: 'We found no evidence for the delayed issue effect; i.e. the effort to resolve issues in a later phase was not consistently or substantially greater than when issues were resolved soon after their introduction.'

The classic exponential cost-of-delay curve does not replicate, so the case for planning before an agent runs has to rest on measured agent failure rates rather than shift-left folklore.

2 more excerpts
  • Requirements issues reaching system test showed roughly a 1.85x median resolution-time increase, against the 37-250x multipliers cited in the classic literature
  • Used in the post as an honesty move - it argues against a convenient cliche the author declined to use

Are Delayed Issues Harder to Resolve? Revisiting Cost-to-Fix of Defects throughout the Lifecycle ↗·Cited in Task Decomposition for AI Coding Agents: Draw the Graph First

Xueping Gao Jun 16, 2026 data

Across CompSkillBench - 300 compositional queries over 2,209 real MCP server skills spanning 24 categories - standard LLM decomposition reaches only 34.2% category recall at the step level, making decomposition quality the primary bottleneck.

Granularity is the hard part of decomposition and the part models are worst at, which is why the task graph is drawn by a human rather than delegated to the agent.

1 more excerpt
  • Iterative Skill-Aware Decomposition raised decomposition accuracy from 51.0% to 67.7% (+32.7%, Wilcoxon p < 10^-6)

Compositional Skill Routing for LLM Agents: Decompose, Retrieve, and Compose ↗·Cited in Task Decomposition for AI Coding Agents: Draw the Graph First

Addy Osmani Jan 28, 2026 practitioner partial

only 48% of developers consistently check AI-assisted code before committing it, even though 38% find that reviewing AI-generated logic actually requires more effort than reviewing human-written code.

Most teams under-review AI code even though reviewing it costs more effort, so the last-mile verification tax is real and often unpaid.

1 more excerpt
  • AI gets you 80% to an MVP; the last 20% requires patience, learning deeply or hiring engineers.

The 80% Problem in Agentic Coding ↗·Cited in What AI Coding Agents Are Actually Good For (And When to Skip)

Addy Osmani January 5, 2026 practitioner partial

AI writes faster. Humans still have to prove it works.

AI speeds up writing but shifts the constraint to verification; a human still owns proving the code works.

3 more excerpts
  • If your pull request doesn't contain evidence that it works, you're not shipping faster
  • 45% of AI-generated code contains security flaws
  • Logic errors appear at 1.75× the rate of human-written code

Code Review in the Age of AI ↗·Cited in What AI Coding Agents Are Actually Good For (And When to Skip)

Anthropic Mar 25, 2026 practitioner

A user asked to "clean up old branches." The agent listed remote branches, constructed a pattern match, and issued a delete. This would be blocked since the request was vague, the action irreversible and destructive, and the user may have only meant to delete local branches.

Vague-plus-irreversible-plus-destructive is the dangerous combination to gate; a concrete incident shows why you don't delegate blast-radius actions blind.

4 more excerpts
  • Claude Code users approve 93% of permission prompts.
  • If a session accumulates 3 consecutive denials or 20 total, we stop the model and escalate to the human.
  • Destroy or exfiltrate. Cause irreversible loss by force-pushing over history, mass-deleting cloud storage, or sending internal data externally.
  • Instead, a false positive costs a single retry where the agent gets a nudge, reconsiders, and usually finds an alternative path.

How we built Claude Code auto mode: a safer way to skip permissions ↗·Cited in What AI Coding Agents Are Actually Good For (And When to Skip), Tier Your AI Agent's Production Authority by Task Risk

Anthropic (Erik Schluntz and Barry Zhang) Dec 19, 2024 practitioner

When building applications with LLMs, we recommend finding the simplest solution possible, and only increasing complexity when needed. This might mean not building agentic systems at all.

The default should be the simplest solution; reaching for an agent is a decision to justify, not an assumption.

4 more excerpts
  • They are typically just LLMs using tools based on environmental feedback in a loop.
  • Code solutions are verifiable through automated tests; Agents can iterate on solutions using test results as feedback
  • The autonomous nature of agents means higher costs, and the potential for compounding errors.
  • Agentic systems often trade latency and cost for better task performance, and you should consider when this tradeoff makes sense.

Building effective agents ↗·Cited in What AI Coding Agents Are Actually Good For (And When to Skip), Loop Engineering Breaks Your Single-Shot Context Playbook

Birgitta Böckeler (martinfowler.com) Oct 15, 2025 practitioner

Spec-driven tooling applied to a small bug produced four user stories with sixteen acceptance criteria - a sledgehammer for a nut. 'All SDD approaches and definitions I've found are spec-first, but not all strive to be spec-anchored or spec-as-source.'

Over-specification is a real failure mode with a cost, which supplies the stop condition in the post's pricing checklist.

1 more excerpt
  • The most authoritative non-vendor treatment of the spec-first / spec-anchored / spec-as-source distinction found in this research pass

Understanding Spec-Driven-Development: Kiro, spec-kit, and Tessl ↗·Cited in Task Decomposition for AI Coding Agents: Draw the Graph First

Daniel Epstein (Microsoft Developer Blog) May 19, 2026 practitioner

Names the failure directly: 'No backlog: There is no structured list of what needs to be built, in what order, with what dependencies. Work gets discovered during implementation, not planned before it.' The prescribed fix is 'Specs in Backlog first: Every capability is an issue. Every issue has acceptance criteria.'

A named practitioner framing of the missing artifact this post builds - the dependency-ordered backlog that precedes any agent run.

2 more excerpts
  • The article contains no numbers, percentages, or named studies - it is argumentative, and the post cites it as practitioner framing only, never as measurement
  • A widely circulated line about 'the hardest step ... assumed rather than solved' is from a reader comment by Rolf Kristensen, not from Epstein's article, and is not cited in this post

Agentic-Agile: Why Agent Development Needs Agile (Not Just Prompts) ↗·Cited in Task Decomposition for AI Coding Agents: Draw the Graph First

Hamel Husain, interviewed by Sara Verdi (Arize) Jul 30, 2026 practitioner

'A really common way that the model is not the problem is query disambiguation. The LLM doesn't have a chance because the user is asking a very ambiguous question.'

Ambiguity attaches to interface contracts, not just prose - an agent told to clean up an authentication service cannot know whether it may change the public API, add a dependency, touch the schema, or alter error behavior.

1 more excerpt
  • Backs the semantic half of the post's two-edge model: a node needs an owned contract, not only an owned file glob

Rise of the Agent Engineer: Why AI Evals Fail Before the Evaluation Begins ↗·Cited in Task Decomposition for AI Coding Agents: Draw the Graph First

Kent Beck (O'Reilly, 'Coding with AI: The End of Software Development As We Know It') session page, undated; underlying event May 8, 2025 practitioner partial

'Augmented coding deprecates formerly leveraged skills such as language expertise. Augmented coding amplifies vision, strategy, task breakdown, and feedback loops.'

Task breakdown is an appreciating skill under agentic coding, not a depreciating one - which is the argument for naming an owner rather than letting it go unassigned.

2 more excerpts
  • Quote confirmed verbatim on the O'Reilly session page, but the page carries no publication date; the event date is corroborated from independent announcements and O'Reilly Radar coverage
  • Beck's own newsletter does not carry this exact sentence; kentbeck.com has a near-identical paraphrase in a mutable homepage section, so the O'Reilly page is the only stable surface for the verbatim wording

Vibe Coding: More Experiments, More Care ↗·Cited in Task Decomposition for AI Coding Agents: Draw the Graph First

Sean Goedecke May 17, 2026 practitioner

For difficult tasks, I'll often reject five or six (or more!) agent attempts before accepting one as good enough to work with, or giving up and making the change by hand.

Getting value from agents on hard tasks means aggressively rejecting weak attempts and keeping judgment work human.

3 more excerpts
  • able to correctly diagnose 80% of issues on its own
  • The current core AI skill is shifting as much work onto AI agents as possible, without going too far.
  • I still don't use LLMs to write Slack messages, ADRs, issues and so forth.

How I use LLMs as a staff engineer in 2026 ↗·Cited in What AI Coding Agents Are Actually Good For (And When to Skip)

Sean Goedecke February 4, 2025 practitioner

LLMs excel at writing code that works that doesn't have to be maintained.

Agents are best on throwaway and research code, not the maintained business logic and judgment writing you own long-term.

3 more excerpts
  • I would say that my use of LLMs here meant I got this done 2x-4x faster
  • It's rare that I let Copilot produce business logic for me
  • I **never** allow the LLM to write these for me

How I use LLMs as a staff engineer ↗·Cited in What AI Coding Agents Are Actually Good For (And When to Skip)

Simon Willison 7th October 2025 practitioner

If your project has a robust, comprehensive and stable test suite agentic coding tools can _fly_ with it.

A strong automated test suite is the single biggest enabler of agent productivity on a codebase.

2 more excerpts
  • what should we call the other end of the spectrum, where seasoned professionals accelerate their work with LLMs while staying proudly and confidently accountable for the software they produce?
  • Automated testing / Planning in advance / Comprehensive documentation / Good version control habits / Effective automation / Culture of code review / Manual QA / Research skills / Ship to preview environment

Vibe engineering ↗·Cited in What AI Coding Agents Are Actually Good For (And When to Skip)

Agent Runtime

132 sources

The machinery around the model — the context it sees, the harness it acts through, the loop it runs in.

Anthropic June 13, 2025 data

A four-field delegation spec for every subagent: an objective, an output format, guidance on the tools and sources to use, and clear task boundaries. Without it, one subagent explored the 2021 automotive chip crisis while two others duplicated work on 2025 supply chains. Agents use about 4x more tokens than chat interactions; multi-agent systems about 15x.

Every edge in an agent graph needs an explicit specification, and the failure mode of an underspecified edge is duplicated and misaligned work rather than a visible error.

4 more excerpts
  • 'Agents make dynamic decisions and are non-deterministic between runs, even with identical prompts.' Scale context: 'Simple fact-finding requires just 1 agent with 3-10 tool calls, direct comparisons might need 2-4 subagents with 10-15 calls each'.
  • Subagents facilitate compression by operating in parallel with their own context windows, exploring different aspects of the question simultaneously before condensing the most important tokens for the lead research agent.
  • Coding-specific bound, time-indexed to June 2025: 'most coding tasks involve fewer truly parallelizable tasks than research, and LLM agents are not yet great at coordinating and delegating to other agents in real time.'
  • The 15x and 4x multiples are both baselined against chat interactions, not against single-agent search - not interchangeable with the 3-10x figure from Anthropic's January 2026 post.

How we built our multi-agent research system ↗·Cited in The Topology You Can Review Is the Topology You Can Run, Rank Your Agent Trace Fields by What You Can Never Get Back, When One Agent Stops Being Enough: The Isolation Gate

Anthropic November 24, 2025 data

The most common failures are wrong tool selection and incorrect parameters, especially when tools have similar names like `notification-send-user` vs. `notification-send-channel`.

At scale, loading all tool definitions upfront is the failure driver; deferred tool loading cuts token cost and measurably raises tool-selection accuracy.

4 more excerpts
  • When using natural language tool calling, each invocation requires a full inference pass, and intermediate results pile up in context whether they're useful or not.
  • At Anthropic, we've seen tool definitions consume 134K tokens before optimization.
  • Opus 4.5 improved from 79.5% to 88.1%
  • This represents an 85% reduction in token usage while maintaining access to your full tool library.

Introducing advanced tool use on the Claude Developer Platform ↗·Cited in Where Just-in-Time Context Retrieval Silently Breaks, Loop Engineering Breaks Your Single-Shot Context Playbook

Anthropic Frontier Red Team August 13, 2026 data

In a shared-repository run, 18 out of 30 agents created a git branch with the identical name 'mvp-game-loop'. A job-market coordination attempt produced 2.4 million job requests against 117 accepted jobs. On hidden-profile tasks, groups scored 17-36% against solo ceilings near 100%. In a coordinated vulnerability hunt, agents found 266 vulnerabilities versus 21 in a solo 6.5-million-token run, with only 12 in common.

Coordination failures in multi-agent systems are properties of the arrangement rather than of model capability, and they show up first in ordinary coding artifacts: colliding branch names, conflicting pull requests, and duplicated work.

3 more excerpts
  • Structure determines behavior directly: given a private back-channel, agents began colluding almost immediately.
  • Correlated failure across a fleet: when one agent makes a bad decision, many agents are likely to make the same one.
  • The words 'checkpoint', 'approval', 'topology', 'graph' and 'architecture' do not appear anywhere in the piece; human intervention appears only as a fallback.

Patterns and problems in emerging multiagent systems ↗·Cited in The Topology You Can Review Is the Topology You Can Run

Bandi, Dumitru, Hertzberg, Agarwal et al. (Scale AI) Jan 31, 2026 data

Across 1,000 expert-written tasks spanning 36 real MCP servers and 220 tools, automated diagnostics show 63.3% of diagnosed failures are cognitive rather than tool-call related.

Second independent finding that the majority of agent failures are model-side, not tool-surface.

1 more excerpt
  • 'Several high-performing models fail after successful tool execution due to premature stopping or incorrect synthesis' - a cognitive fault the harness can still address

MCP-Atlas: A Large-Scale Benchmark for Tool-Use Competency with Real MCP Servers ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Charles Fleming, Guillaume De Saint Marc, Ramana Kompella, Peter Bosch, Vijoy Pandey v1 Apr 3, 2026; v2 May 27, 2026 data

A service-mesh-style middleware layer attached to every agent instance improved performance more than 10% over direct agent-to-agent communication on HotPotQA (80.1 to 91.5 vs. 92 baseline) and MuSiQue (72.7 to 86.1 vs. 87.5 baseline).

Direct, prompt-embedded agent-to-agent communication measurably costs accuracy; a separated coordination layer recovers most of it.

1 more excerpt
  • The PDF returned only partial text on direct fetch; verified via the HTML render and cross-checked against the abstract's summary figure

Scaling Multi-agent Systems: A Smart Middleware for Improving Agent Interactions ↗·Cited in Coordination Is an Architecture Layer, Not a Prompt Instruction

Daniel Jaroslawicz, Brendan Whiting, Parth Shah, Karime Maamari (Distyl AI) 2025 (arXiv 2507.11538v1) data

Even the best frontier models only achieve 68% accuracy at the max density of 500 instructions.

Instruction-following accuracy degrades sharply with density — the best frontier models hit only 68% at 500 instructions — so packing rules in measurably erodes compliance.

4 more excerpts
  • At 500 instructions, llama-4-scout exhibits an extreme O:M ratio of 34.88, indicating omission errors are over 30 times more frequent
  • Threshold decay: "Performance remains stable until a threshold, then transitions to a different (steeper) degradation slope" — exhibited by gemini-2.5-pro, o3
  • Primacy effects display an interesting pattern across all models: they start low at minimal instruction densities indicating almost no bias for earlier instructions, peak around 150–200 instructions
  • Analysis reveals model size and reasoning capability to correlate with 3 distinct performance degradation patterns, bias towards earlier instructions, and distinct categories of instruction-following errors.

How Many Instructions Can LLMs Follow at Once? ↗·Cited in CLAUDE.md Instruction Ceiling: Maintained Config, Not a README, How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

Dat Tran, Douwe Kiela April 2, 2026 (v2: April 11, 2026) data

Across two datasets (FRAMES, MuSiQue), three model families (Qwen3, DeepSeek, Gemini) and five multi-agent architectures (Sequential, Debate, Ensemble, Parallel-roles, Subtask-parallel), single-agent systems match or outperform multi-agent systems when reasoning tokens are held constant. Thinking-token budgets tested at 100, 500, 1k, 2k, 5k and 10k.

Many reported multi-agent advantages are better explained by unaccounted computation and context effects than by any inherent architectural benefit, so an edge's apparent gain has to be re-measured with compute held equal before it is kept.

3 more excerpts
  • The paper carves out a regime where multi-agent wins: in sufficiently degraded-context regimes, structured pipelines 'may occasionally surpass SAS by imposing useful factorization, filtering, or verification structure'.
  • Explicitly scoped out of tool-using settings: 'We focus on text-only multi-hop reasoning; MAS advantages with tools/vision or safety constraints are out of scope.' Coding agents are tool-using, so this is indirect evidence for coding.
  • Non-peer-reviewed preprint.

Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets ↗·Cited in The Topology You Can Review Is the Topology You Can Run

Dun Yuan, Fuyuan Lyu, Ye Yuan, Weixu Zhang, Bowei He, Jiayi Geng, Linfeng Du, Zipeng Sun, Yankai Chen, Changjiang Han, Jikun Kang, Xi Chen, Haolun Wu, Xue Liu v1 Mar 30, 2026; v3 Apr 13, 2026 data

A systematic audit of 18 agent communication protocols found semantic-layer mechanisms (clarification, context alignment, verification) largely absent at the protocol level.

Absent a protocol-level home for coordination logic, developers reintroduce it through prompts, wrappers, and orchestration glue by default, which accumulates hidden complexity.

1 more excerpt
  • Evidence base is a qualitative comparative audit across nine dimensions in three layers, not quantitative failure-rate measurement

Beyond Message Passing: A Semantic View of Agent Communication Protocols ↗·Cited in Coordination Is an Architecture Layer, Not a Prompt Instruction

Gloaguen, Mündler, Müller, Raychev, Vechev (ETH Zurich) Feb 12, 2026 (revised Jun 23, 2026) data

Providing context files does not generally improve task success rates while increasing inference cost by over 20% on average. Developer-provided files improved performance by 2.4% on average, not statistically significant (p=21%); LLM-generated files caused drops in 5 of 8 settings.

The most portable layer of the setup is not the most valuable one, so portability of AGENTS.md should not be mistaken for leverage.

4 more excerpts
  • we find that context files tend to reduce task success rates compared to providing no repository context, while also increasing inference cost by over 20%.
  • Context files increased steps in every setting, by 2.45 and 3.92 on average, driving cost increases of 20% and 23% on SWE-bench and CTXbench respectively
  • Developer-provided files significantly outperform LLM-generated ones (p=3.8%)
  • Instructions are well followed; repository overviews, although recommended by model providers, are not helpful

Evaluating AGENTS.md: Are Repository-Level Context Files Helpful for Coding Agents? ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness, How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

Huang et al. (arXiv) Submitted 22 January 2026 (accepted at ICAIBD 2026) data partial

Procedural reliability, particularly tool initialization failures, constitutes the primary bottleneck for smaller models.

For smaller models, tool-invocation reliability (especially tool initialization) is the primary failure bottleneck, localizable via a 12-category taxonomy.

3 more excerpts
  • 1,980 deterministic test instances
  • 12-category error taxonomy capturing failure modes across tool initialization, parameter handling, execution, and result interpretation
  • Mid-sized models (qwen2.5:14b) offer practical accuracy-efficiency trade-offs on commodity hardware (96.6% success rate, 7.3 s latency)

When Agents Fail to Act: A Diagnostic Framework for Tool Invocation Reliability in Multi-Agent LLM Systems ↗·Cited in Where Just-in-Time Context Retrieval Silently Breaks

Kelly Hong, Anton Troynikov, Jeff Huber (Chroma) July 14, 2025 data

Even under these minimal conditions, model performance degrades as input length increases, often in surprising and non-uniform ways.

18 LLMs degrade non-uniformly as input grows — the independent mechanism behind the bloat warning (applies to CLAUDE.md by analogy; the study never tests it).

4 more excerpts
  • models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows
  • Even a single distractor reduces performance relative to the baseline (needle only).
  • Whether relevant information is present in a model's context is not all that matters; what matters more is how that information is presented.
  • Even a single distractor reduces performance relative to the baseline (needle only), and adding four distractors compounds this degradation further

Context Rot: How Increasing Input Tokens Impacts LLM Performance ↗·Cited in How to Structure CLAUDE.md: It's a Loading Policy, Not a Document, CLAUDE.md Instruction Ceiling: Maintained Config, Not a README, Loop Engineering Breaks Your Single-Shot Context Playbook, Where Just-in-Time Context Retrieval Silently Breaks, When One Agent Stops Being Enough: The Isolation Gate

Liu et al. (TACL 2024) July 6, 2023 (v1); revised Nov 20, 2023 data

language model performance is highest when relevant information occurs at the very beginning (primacy bias) or end of its input context (recency bias), and performance significantly degrades when models must access and use information in the middle of long contexts

Mid-file placement is the worst-served position — the U-shaped retrieval curve.

1 more excerpt
  • extended-context models are not necessarily better than their non-extended counterparts at using their input context.

Lost in the Middle: How Language Models Use Long Contexts ↗·Cited in How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

Matthias Galster, Seyedmoein Mohsenimofidi, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian Baltes Feb 16, 2026 (v5: Jun 30, 2026) data

Across 2,853 repositories, 2,015 (70.6%) adopted a single tool while 295 (10.3%) configured two, and 493 repositories (17.3%) used AGENTS.md as a tool-agnostic standard without any tool-specific artifact. The paper recommends that developers who rely on multiple tools maintain an AGENTS.md file as the shared core configuration, with tool-specific files as adapters that reference it.

The shared-core-plus-adapters layout is the empirically observed pattern and the paper's own recommendation, not a preference of the post.

4 more excerpts
  • Keep repo-level adoption (AGENTS.md 39.5%, CLAUDE.md 45.9%) separate from file-count share (31.6% and 34.4%); the paper reports both
  • No mechanism beyond context files exceeds 20% adoption for Claude, Copilot, Cursor or Gemini; 72.8% of Cursor repositories adopt Rules
  • Adopting tool-specific mechanisms such as Rules or Commands ties workflows to a specific tool, in the paper's own words
  • 85.5% of Skills include no additional resources, so Skills function primarily as structured text

Harness Engineering for Agentic AI Coding Tools: An Exploratory Study ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Paul Barbaste, Tristan Darrigol, Germain Vu, Tom Wiltberger Jul 15, 2026 (expanded second edition of an April 2026 study) data partial

Across eleven harnesses and roughly four million lines of source, SKILL.md skills lead MCP in adoption (9/11 vs 8/11), OpenHands runs Claude Code, Codex, or Gemini CLI as interchangeable backends, and Codex adopts Claude Code's hook vocabulary verbatim and ships an importer for its sessions and settings.

The harness layer has converged enough that one harness can host its rivals, and behavioral policy is migrating from prompt prose into configuration.

3 more excerpts
  • Per-harness hook event counts (Gemini CLI 11, Pi about 33, OpenCode 20), the 'hooks in 9/11' aggregate and the Polly per-task worktree example could not be located in the paper; those figures are sourced first-party from vendor docs instead
  • No agent runtime in the corpus imports a general-purpose agentic framework or retrieves code with vector embeddings
  • Preprint; bibliography was not extractable in the deep read

Harness Engineering: Anatomy, Architecture, and Evolution of Coding Agents, A Source-Code Study of Eleven Systems ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Ponnusamy, Chandran, Hossain 25 Dec 2025 data

For Llama-3.1-70B, accuracy declined only slightly from the 98.5% baseline to 98% at 15,000 words.

Truly unrelated filler barely dents accuracy — the honest qualifier on the distractor analogy; the bill is latency, not correctness.

1 more excerpt
  • the observed 719.64% increase in latency for the 70B model at the 15,000-word regime

Context Discipline and Performance Correlation: Analyzing LLM Performance and Quality Degradation Under Varying Context Lengths ↗·Cited in How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

Shi et al. (ICML 2023) 31 Jan 2023 (v1) data

a single piece of irrelevant information can distract the models and substantially degrade their performance, even on problems whose clean versions they correctly solve.

Irrelevant context degrades accuracy even when all relevant information is present.

1 more excerpt
  • we find that simply adding an instruction to ignore irrelevant information brings notable performance gains on our benchmark.

Large Language Models Can Be Easily Distracted by Irrelevant Context ↗·Cited in How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

Terminal-Bench (Stanford x Laude) live leaderboard (fetched Jul 27, 2026) data

Eleven Claude Opus 4.6 entries span 58.0% (Claude Code, +/-2.9) to 76.4% (Meta-Harness, +/-2.4) on 89 terminal tasks - an 18.4-point spread on identical model weights, 13.1 points at the non-overlapping confidence bounds.

The same frontier model varies by double-digit percentage points purely as a function of the harness it runs in.

2 more excerpts
  • Entries are self-submitted by harness authors via pull request, machine-validated for timeout/resource parity and a five-trial minimum, then maintainer-merged
  • Submissions span Dec 2025 to May 2026 and are not contemporaneous; harness-side and model-side settings are not held constant, so the spread is observational rather than controlled

Terminal-Bench 2.0 Leaderboard ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

The Terminal-Bench Team (Kelly Buchanan, TB2.1 Lead) data partial

On Terminal-Bench 2.1 the same model scores differently under different harnesses: Opus 4.6 at 70.1% in Claude Code vs 63.8% in Terminus 2, GPT-5.4 at 77.3% in Codex CLI vs 54.8% in Terminus 2, Gemini 3.1 Pro at 70.7% in Terminus 2 vs 67.1% in Gemini CLI.

A model-versus-model leaderboard cannot tell you how a swap will perform, because the harness moves the score.

3 more excerpts
  • The page never states a harness-versus-harness comparison in prose; the pairs are read off the agent-model table, and the harness-causation argument is the post's inference
  • The release fixes 28 of the 89 tasks in Terminal-Bench 2.0; the largest gain is Claude Code with Opus 4.6, up 12.1%
  • Additional pairs: GPT-5.4 mini 66.1% in Codex CLI vs 36.9% in Terminus 2; Sonnet 4.6 58.5% in Claude Code vs 51.5%

Terminal-Bench 2.1 ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Worawalan Chatlatanagulchai et al. 17 Nov 2025 (submitted) data

While developers use context files to make agents functional, they provide few guardrails to ensure that agent-written code is secure or performant

Empirically, teams pack context files with functional setup but almost no security or performance guardrails — the constraint side of CLAUDE.md is systematically under-specified.

3 more excerpts
  • 2,303 agent context files across 1,925 repositories
  • Build and run commands: 62.3%, Implementation details: 69.9%, Architecture: 67.7%; Security: 14.5%, Performance: 14.5%
  • These files are not static documentation but complex, difficult-to-read artifacts that evolve like configuration code

Agent READMEs: An Empirical Study of Context Files for Agentic Coding ↗·Cited in CLAUDE.md Instruction Ceiling: Maintained Config, Not a README

Xiaoyu Chu, Sacheendra Talluri, Qingxian Lu, Alexandru Iosup (VU Amsterdam) Jan 21, 2025 (v2: Mar 15, 2025; ICPE '25) data

For Anthropic's services, the likelihood of any two services experiencing outages on the same day is over 80%, while no correlation is observed between services from different providers. Anthropic averaged an MTTR of 2.70 hours and an MTBF of 5.22 days across the study window.

A vendor's surfaces fail together, so the vendor, not the individual product, is the failure domain to plan around.

4 more excerpts
  • Data window ends 2024-08-31 across 8 services from OpenAI, Anthropic and Character.AI; the over-80% figure is Anthropic-specific
  • The mechanism is hedged: the paper says the difference may be caused by different cloud infrastructures (OpenAI on Azure, Anthropic on GCP)
  • The 49.21% OpenAI figure is API-to-ChatGPT co-occurrence specifically, not a whole-provider aggregate
  • Only 6.15% of reports disclose a postmortem; Anthropic provided none for its API and Console services

An Empirical Characterization of Outages and Incidents in Public Services for Large Language Models ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Yang et al. (NeurIPS 2024) May 6, 2024 (v3 November 11, 2024) data partial

SWE-agent solves 10.7 percentage points more issues than the baseline agent that uses just the default Linux shell (300-issue ablation); on SWE-bench Lite, SWE-agent with GPT-4 Turbo resolves 18.00% versus 11.00% for the shell-only agent with the same model.

Agent-computer interface design alone moves resolve rates by double digits with the model held fixed - the founding demonstration of the scaffold confound.

2 more excerpts
  • Verification is partial because arXiv HTML routes 404 and the PDF required reader-proxy extraction; figures converged across three independent extraction passes
  • Citation caution: the paper's verbatim 10.7pp sentence arithmetically pairs with the 7.33% no-demonstration shell baseline, not the 11.00% row - quote the sentence or the 18.00/11.00 pair, never 10.7 with 11.00

SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Yubin Kim, Ken Gu, Chanwoo Park, Chunjong Park, Samuel Schmidgall, et al. (Google Research / MIT Media Lab / DeepMind) December 9, 2025 (v3: April 8, 2026) data

260 configurations across six agentic benchmarks, five canonical architectures (Single-Agent plus Independent, Centralized, Decentralized, Hybrid) and three LLM families, with tools, prompts and compute standardized to isolate architectural effects. On SWE-bench Verified every multi-agent architecture underperformed the single-agent baseline (mean 0.522): Hybrid -2.1%, Centralized -3.1%, Decentralized -5.4%, Independent -14.9%, on 20-instance subsets. Agent count was not a significant predictor (log(1+n_a): beta=0.040, 95% CI [-0.074, 0.155], p=0.487). Capability ceiling beta=-0.236, p=0.004. Cross-validated R^2=0.373.

The shape of an agent system, not the number of agents in it, determines whether collaboration helps; and on coding work specifically every multi-agent arrangement tested lost to a single agent.

4 more excerpts
  • Relative performance change compared to single-agent baseline ranges from +80.8% on decomposable financial reasoning to -70.0% on sequential planning, demonstrating that architecture-task alignment determines collaborative success.
  • The headline range quoted from the abstract (+80.8% to -70.0%) spans two different architectures on two different tasks: +80.8% is Centralized on Finance Agent, -70.0% is Independent on PlanCraft. Holding Centralized fixed, the pair is +80.8% / -50.3%.
  • Architectures without centralized verification propagate errors more than centrally coordinated ones.
  • Cite arXiv v3; Google's blog post describes an earlier 180-configuration, four-benchmark version.

Towards a Science of Scaling Agent Systems ↗·Cited in The Topology You Can Review Is the Topology You Can Run, When One Agent Stops Being Enough: The Isolation Gate

Zeng et al. (arXiv) Submitted 26 May 2026 data

We attribute this improvement to the legibility of failed logical search. Repeated failures under explicit lexical constraints provide a clearer signal that required evidence may be absent, whereas Agentic Hybrid may still return semantically related but unsupported passages.

Logical/lexical retrieval can signal 'nothing found' where embedding search cannot, which measurably reduces hallucination on answer-unavailable questions.

3 more excerpts
  • On average, its refusal rate increased from 0.767 to 0.828, while the hallucination rate decreased from 0.128 to 0.083.
  • anchoring the retrieval process in logical queries substantially reduces hallucinations in generated responses.
  • matches a strong agentic hybrid baseline, while substantially reducing construction and serving cost

Rethinking Agentic RAG: Toward LLM-Driven Logical Retrieval Beyond Embeddings ↗·Cited in Where Just-in-Time Context Retrieval Silently Breaks

Zhang et al. May 7, 2026 data

In a controlled 3x3 factorial experiment, average harness variance is 18.48 pp-squared versus average model variance of 2.37 pp-squared - a 7.80x ratio - and public leaderboards show harness-only swings of 7.3pp (Terminal-Bench 2), 9.5pp (SWE-bench Pro, same Opus 4.5), up to 15pp (SWE-bench Verified), and 34-48pp cross-scaffold gaps on the HAL Leaderboard.

Performance variance is governed more by harness configuration than model choice, so evaluation protocols without harness disclosure systematically misattribute harness gains to model improvements.

2 more excerpts
  • The same model under a different harness can rank above or below a competitor - rank order itself is harness-dependent
  • The paper proposes a harness-aware evaluation framework with a disclosure standard and variance decomposition protocol

Stop Comparing LLM Agents Without Disclosing the Harness ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Zhao, Li, Li, Zhao, Barr, Sarro, Ye Jul 10, 2026 data

Across 1,794 manually annotated trajectories (63,000+ execution steps, seven frontier models, three scaffolds), environment triggers account for 9.4% of decisive errors against 57.9% epistemic and 32.8% competence.

The direct refutation of harness causation - agent failures are predominantly model-side, which is why this post argues leverage rather than cause.

3 more excerpts
  • Largest single trigger is false premises at 30.7%
  • Epistemic errors are the largest share in every scaffold tested, ranging 44% to 80%
  • The paper's own prescription - 'earlier validation and intervention' - is itself a harness prescription

Failure as a Process: An Anatomy of CLI Coding Agent Trajectories ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Addy Osmani March 26, 2026 practitioner

A six-step production line (Plan, Spawn, Monitor, Verify, Integrate, Retro) with three to five teammates named as the sweet spot and a ratio of one reviewer per three to four builders. Guardrail defaults include MAX_ITERATIONS=8 and an auto-pause at 85% of budget.

Verification, not generation, is the bottleneck in multi-agent coding, and human review is infrastructure rather than optional overhead.

4 more excerpts
  • Three to five teammates is the sweet spot. Token costs scale linearly, and three focused teammates consistently outperform five scattered ones.
  • Review is positioned as a stage in an ordered pipeline, not as a node in a topology; team size is chosen first for parallelism and cost, and the reviewer ratio is applied afterward.
  • Cites Gloaguen et al. (ETH Zurich) for a roughly 3% success-rate reduction and over 20% inference-cost increase from LLM-generated AGENTS.md files.
  • The bottleneck is no longer generation. It's verification.

The Code Agent Orchestra - what makes multi-agent coding work ↗·Cited in The Topology You Can Review Is the Topology You Can Run, When One Agent Stops Being Enough: The Isolation Gate

Addy Osmani June 7, 2026 practitioner

Loop engineering is replacing yourself as the person who prompts the agent. You design the system that does it instead.

Loop engineering shifts leverage from writing prompts to designing the loop that prompts the agent.

1 more excerpt
  • The hard part is not autonomy itself. It is verification, stopping conditions, and Human in the Loop escalation.

Loop Engineering ↗·Cited in Loop Engineering Breaks Your Single-Shot Context Playbook

Addy Osmani May 3, 2026 practitioner

'The same SKILL.md file works in Claude Code, Cursor (with rules), Gemini CLI, Codex, and any other harness that accepts system-prompt content.' Cursor users put them in .cursor/rules/; Gemini CLI has its own install path.

Skills are the one advanced mechanism practitioners report carrying across harnesses, because they are plain markdown with frontmatter.

2 more excerpts
  • Slash commands sit on top as an add-on layer (seven commands over twenty skills in the author's repo)
  • Practitioner report on one author's repo, which had crossed 27K stars; not a controlled test of portability

Agent Skills ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Agent Skills (agentskills.io) practitioner

Only name (max 64 characters, lowercase, must match the directory) and description (max 1024 characters) are required; six fields total are in the frontmatter table, and allowed-tools is 'Experimental. Support for this field may vary between agent implementations.'

The portable surface of a skill is the six-field spec, and anything beyond it is per-tool.

3 more excerpts
  • The spec does not state how agents must treat unrecognized top-level frontmatter fields; metadata is the sanctioned place for non-spec data
  • Progressive disclosure: about 100 tokens of metadata loaded at startup, a body under 5,000 tokens recommended, main SKILL.md under 500 lines
  • Script runtimes depend on the agent implementation

Specification ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Agentic AI Foundation (Linux Foundation) undated (fetched Jul 13, 2026) practitioner partial

Over 60,000 open-source projects use AGENTS.md; 'the closest AGENTS.md to the edited file wins; explicit user chat prompts override everything.' The supported-tools list runs to 23 entries including OpenAI Codex, Cursor and Google Gemini CLI, and Claude Code is not on it.

AGENTS.md is the vendor-neutral instruction convention, and Claude Code's absence from the list is why the post bridges it with an import.

4 more excerpts
  • Agents automatically read the nearest file in the directory tree, so the closest one takes precedence and every subproject can ship tailored instructions.
  • Some strings came back as the fetch tool's paraphrase; re-verify exact wording before quoting verbatim
  • Per-tool flips shown on the page: Aider via .aider.conf.yml read: AGENTS.md, Gemini CLI via .gemini/settings.json context filename
  • Stewarded by the Agentic AI Foundation under the Linux Foundation

AGENTS.md ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness, How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

Aider (Paul Gauthier) continuously updated (fetched August 2026) practitioner

The same model's code-editing score moves ~10 points by edit format alone: gemini-exp-1206 scores 80.5% in whole format versus 69.2% in diff format; o1-mini 70.7% versus 61.1%.

A live, reproducible public leaderboard shows the harness's edit protocol moving scores by roughly the same magnitude as top-of-table model gaps.

1 more excerpt
  • Live page - re-pull current figures before quoting in new work

Code editing leaderboard ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Albert Nahas Feb 17 (year not stated on page; brief lists 2026) practitioner

when the context window fills up and gets compacted, your CLAUDE.md values get summarized away with everything else

CLAUDE.md instructions decay mid-session — they get summarized away at compaction — so hook-based reinforcement is more reliable for must-follow standards.

3 more excerpts
  • hook output requires approximately 15 tokens per prompt reminder
  • Over 50-turn session, motto reminders total ~750 tokens against 200k context window
  • hook output arrives as clean system-reminder messages — no disclaimer, no 'may or may not be relevant' framing

Your CLAUDE.md Instructions Are Being Ignored - Here's Why (and How to Fix It) ↗·Cited in CLAUDE.md Instruction Ceiling: Maintained Config, Not a README

Anthropic September 29, 2025 practitioner

Of course, there's a trade-off: runtime exploration is slower than retrieving pre-computed data. Not only that, but opinionated and thoughtful engineering is required to ensure that an LLM has the right tools and heuristics for effectively navigating its information landscape.

Just-in-time context retrieval is not free: it trades latency for freshness and demands deliberate tool and heuristic design to work.

4 more excerpts
  • An agent running in a loop generates more and more data that could be relevant for the next turn of inference, and this information must be cyclically refined.
  • Context, therefore, must be treated as a finite resource with diminishing marginal returns.
  • agents built with the 'just in time' approach maintain lightweight identifiers (file paths, stored queries, web links, etc.) and use these references to dynamically load data into context at runtime using tools.
  • In certain settings, the most effective agents might employ a hybrid strategy, retrieving some data up front for speed, and pursuing further autonomous exploration at its discretion.

Effective context engineering for AI agents ↗·Cited in Where Just-in-Time Context Retrieval Silently Breaks, Loop Engineering Breaks Your Single-Shot Context Playbook, How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

Anthropic 2026 practitioner

As token count grows, accuracy and recall degrade, a phenomenon known as context rot. This makes curating what's in context just as important as how much space is available.

Curating what's in the context window matters as much as how much space is available - official-docs corroboration of context rot.

Context windows ↗·Cited in Loop Engineering Breaks Your Single-Shot Context Playbook

Anthropic Nov 26, 2025 practitioner

The core challenge of long-running agents is that they must work in discrete sessions, and each new session begins with no memory of what came before.

A high-level prompt alone fails a long-running loop; cross-session state must be externalized to disk.

2 more excerpts
  • even a frontier coding model like Opus 4.5 running on the Claude Agent SDK in a loop across multiple context windows will fall short of building a production-quality web app if it's only given a high-level prompt
  • memoryless-session framing is softened by Opus 4.5+ auto-compaction per Anthropic's March 2026 follow-up - the externalized-state lesson persists, the mechanism is version-dependent

Effective harnesses for long-running agents ↗·Cited in Loop Engineering Breaks Your Single-Shot Context Playbook

Anthropic Sept 29, 2025 practitioner

gather context -> take action -> verify work -> repeat

The agent loop is a repeated four-step cycle; managing context across iterations (compaction) is a loop-only concern with no single-task analog.

1 more excerpt
  • The Claude Agent SDK's compact feature automatically summarizes previous messages when the context limit approaches, so your agent won't run out of context.

Building agents with the Claude Agent SDK ↗·Cited in Loop Engineering Breaks Your Single-Shot Context Playbook

Anthropic Oct 16, 2025 practitioner partial

This metadata is the first level of progressive disclosure: it provides just enough information for Claude to know when each skill should be used without loading all of it into context.

Progressive-disclosure mechanics: metadata triggers, bodies load on relevance.

1 more excerpt
  • This means that the amount of context that can be bundled into a skill is effectively unbounded.

Equipping agents for the real world with Agent Skills ↗·Cited in How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

Anthropic April 10, 2026 practitioner

Five named coordination patterns: generator-verifier, orchestrator-subagent, agent teams, message bus, and shared-state. Pattern choice is derived from workflow predictability, decomposability, interdependence and context duration.

The current frontier-lab state of the art for choosing an agent arrangement is a pattern catalogue selected by task structure, with an explicit recommendation to start with the simplest pattern and evolve from there.

3 more excerpts
  • The words 'risk', 'blast radius', 'reviewer', 'review capacity', 'bandwidth' and 'checkpoint' appear nowhere in the article.
  • Human involvement appears exactly once, as a fallback inside a loop-prevention mechanism.
  • Contains no quantitative data of any kind.

Multi-agent coordination patterns: Five approaches and when to use them ↗·Cited in The Topology You Can Review Is the Topology You Can Run

Anthropic Sep 17, 2025 practitioner

Approximately 30% of Claude Code users who made requests during the affected window had at least one message routed to the wrong server type. Misrouting peaked at 16% of Sonnet 4 requests on the first-party platform on August 31, 0.18% on Amazon Bedrock, and under 0.0004% on Google Cloud's Vertex AI.

The same model degraded very differently by serving path, and internal evals did not catch what users were reporting.

4 more excerpts
  • The 30% figure counts users with at least one misrouted message, not sustained degradation, and is scoped to Claude Code users
  • Routing was sticky, so a request served by the wrong server made follow-ups likely to hit it too
  • The piece shows containment within Anthropic's own infrastructure; it cannot speak to other vendors being unaffected in the same window
  • Bug 2 and Bug 3 date boundaries are less crisp in the text; verify before quoting bug-by-bug ranges

A postmortem of three recent issues ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Anthropic (Claude Code docs) undated (fetched Jul 13, 2026) practitioner

Unlike CLAUDE.md content, a skill's body loads only when it's used, so long reference material costs almost nothing until you need it.

Skills are the designated destination for procedures that outgrew CLAUDE.md — with a stickiness caveat once invoked.

4 more excerpts
  • 'Outside Claude Code, you can use only the fields in the Agent Skills spec. If you include any field the spec doesn't allow, packaging or upload fails with a hard error instead of ignoring the field.' Six fields are portable: name, description, license, compatibility, metadata, allowed-tools.
  • When you or Claude invoke a skill, the rendered SKILL.md content enters the conversation as a single message and stays there for the rest of the session.
  • The hard-error enforcement is documented for Anthropic's own paths (claude.ai uploads, the Skills API, package_skill.py); the page never mentions Codex, Cursor or Gemini CLI, so do not cite it as evidence that third-party agents reject non-spec fields
  • Custom commands have been merged into skills: .claude/commands/deploy.md and .claude/skills/deploy/SKILL.md both create /deploy

Extend Claude with skills ↗·Cited in How to Structure CLAUDE.md: It's a Loading Policy, Not a Document, AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Anthropic (Claude Code docs) undated (fetched Jul 27, 2026) practitioner

'They provide deterministic control over Claude Code's behavior, ensuring certain actions always happen rather than relying on the LLM to choose to run them.'

'Deterministic' is the vendor's own word for the control surface - but the claim is immediately qualified.

1 more excerpt
  • The very next sentence: prompt-based and agent-based hooks 'use a Claude model to evaluate conditions' - so 'hook' is an umbrella containing probabilistic members, and the determinism attaches to command hooks only

Automate actions with hooks ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Anthropic (Claude Code docs) undated (fetched Jul 27, 2026) practitioner

'Claude Code only has the permissions you grant it. You're responsible for reviewing proposed code and commands for safety before approval.'

The vendor explicitly transfers review responsibility to the user - the responsibility this post argues is being discharged on the wrong layer.

3 more excerpts
  • 'Fail-closed matching: Unmatched commands default to requiring manual approval' - scoped to bash permission-rule matching, NOT hook failure handling
  • 'Trust verification is disabled when running non-interactively with the -p flag'
  • Safeguard list: network request approval, isolated context windows, trust verification, command injection detection, fail-closed matching, natural language descriptions, secure credential storage

Security ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Anthropic (Claude Code Docs) 2026 (undated on page; brief dates it 2026) practitioner

If Claude keeps doing something you don't want despite having a rule against it, the file is probably too long and the rule is getting lost. If Claude asks you questions that are answered in CLAUDE.md, the phrasing might be ambiguous. Treat CLAUDE.md like code: review it when things go wrong, prune it regularly, and test changes by observing whether Claude's behavior actually shifts.

Anthropic's own guidance says to maintain CLAUDE.md like code — prune it, and test rule changes by observing whether Claude's behavior actually shifts.

4 more excerpts
  • Claude stops when the work looks done. Without a check it can run, 'looks done' is the only signal available, and you become the verification loop: every mistake waits for you to notice it.
  • The over-specified CLAUDE.md. If your CLAUDE.md is too long, Claude ignores half of it because important rules get lost in the noise.
  • Keep it concise. For each line, ask: 'Would removing this cause Claude to make mistakes?' If not, cut it. Bloated CLAUDE.md files cause Claude to ignore your actual instructions!
  • Ruthlessly prune. If Claude already does something correctly without the instruction, delete it or convert it to a hook.

Best practices for Claude Code ↗·Cited in CLAUDE.md Instruction Ceiling: Maintained Config, Not a README, How to Verify AI Coding Agent Output: A Reviewer's Framework, How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

Anthropic (David Dworken, Oliver Weller-Davies) Oct 20, 2025 practitioner

'Constantly clicking approve slows down development cycles and can lead to approval fatigue, where users might not pay close attention to what they're approving, and in turn making development less safe.'

Why prompt-time human review does not scale, argued by the vendor whose product depends on it.

2 more excerpts
  • 'In our internal usage, we've found that sandboxing safely reduces permission prompts by 84%' - vendor-internal, no methodology, sample size, or definition of 'safely'
  • Context-reversal finding: this blog says a successful prompt injection is 'fully isolated', which the product documentation explicitly contradicts ('not a complete isolation boundary')

Beyond permission prompts: making Claude Code more secure and autonomous ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Anthropic (platform docs) undated (fetched Jul 13, 2026) practitioner

Level 1: Metadata | Always (at startup) | ~100 tokens per Skill ... Level 2: Instructions | When Skill is triggered | Under 5k tokens ... Level 3+: Resources | As needed | Effectively unlimited

The on-demand tier has a documented, quantified cost model (Anthropic's stated architecture, not a measured benchmark).

1 more excerpt
  • This filesystem-based architecture enables progressive disclosure: Claude loads information in stages as needed, rather than consuming context upfront.

Agent Skills ↗·Cited in How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

Anthropic (status.claude.com) Sep 3, 2026 practitioner partial

A major-impact incident on 2026-09-03 ran from 13:26 to 16:23 UTC (2h57m) and listed claude.ai, the Claude API, Claude Code and Claude Cowork as affected components.

A model-layer outage takes Claude Code down with every other Anthropic surface at once, which is the drill the post runs.

2 more excerpts
  • No dedicated dossier deep read; the timeline and affected components come from the status.claude.com incidents feed captured on 2026-09-07
  • The incident sits in a rolling status window and will age out of the public feed

Elevated errors for multiple models ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Anthropic (status.claude.com) Jul 17, 2026 practitioner

A single model incident on 2026-07-17 affected claude.ai, Claude API (api.anthropic.com), Claude Code, and Claude Cowork from 06:47 to 12:21 UTC, roughly 5 hours 34 minutes, and the status regressed from Monitoring back to Identified at 07:10 UTC before resolution.

Outages co-occur inside a provider: one model incident took the coding agent down with the API and the chat product.

2 more excerpts
  • Single-vendor incident; says nothing about cross-vendor correlation
  • Non-monotonic status (a fix declared at 07:03 UTC, then reverted) is why a status page reading needs the full timeline, not the latest update

Elevated errors on Sonnet 5 and Haiku 4.5 ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Anthropic (status.claude.com) Sep 9, 2025 to Sep 17, 2025 practitioner partial

The public incident opened on Sep 9, 2025 for a Claude Sonnet 4 bug that started Aug 5 and a second bug affecting Claude Haiku 3.5 and Sonnet 4 from Aug 26; the detection method is recorded as community reports helping identify and isolate the bugs.

Quality degradation ran for weeks before the vendor opened an incident, and users detected it first.

3 more excerpts
  • Affected components list claude.ai, Claude Console, Claude API and Claude Code; single-vendor and cannot speak to cross-vendor correlation
  • A Sep 12 monitoring update said no ongoing issues five days before the Sep 17 resolution
  • The engineering postmortem is the fuller account of the same bugs

Model output quality ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Anthropic (status.claude.com) Aug 28, 2026 practitioner partial

Anthropic's Identified update at 17:22 UTC on 2026-08-28 reads: 'We have identified an issue with an upstream cloud provider affecting Claude Cowork and Claude Code on the web.' The major-impact incident ran until 20:21 UTC, about 2h59m.

An upstream cloud provider, not a model, can take the coding agent down, so the failure domain includes the vendor's own infrastructure dependencies.

3 more excerpts
  • No dedicated dossier deep read; the quote and timeline come from the status.claude.com incidents feed captured on 2026-09-07
  • The upstream provider is not named
  • Scoped to Claude Code on the web and Claude Cowork, not the terminal client

Elevated errors on Claude Code and Claude Cowork ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Anthropic (status.claude.com) undated (live; feed updated 2026-09-07) practitioner partial

As displayed on 2026-09-07, 90-day uptime read 99.44% for Claude Code, 99.5% for the Claude API, 99.4% for claude.ai and 99.43% for Claude Cowork. The incidents feed held 28 entries from 2026-08-04 to 2026-09-03, 19 of them listing Claude Code among affected components.

Anthropic's surfaces fail together repeatedly, and any incident count drawn from the page must carry a date because the feed is a rolling window.

3 more excerpts
  • A prior-pass count (37 of 50 incidents from 2026-07-21) could not be reproduced; the feed's earliest entry had moved to 2026-08-04, so counts shrink over time
  • The 19-of-28 and 9 major-or-critical figures are derived by itemizing the feed, not stated by the page
  • Provides nothing on cross-vendor correlation

Claude Status ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Antigravity CLI docs (antigravity.google) practitioner partial

'Both CLI platforms utilize identical workspace context rules. No modifications are needed to your existing rule documents': the agent continues to parse GEMINI.md and AGENTS.md in the active directory and ~/.gemini/GEMINI.md globally.

The instruction-file layer survived a vendor retiring its own tool; a team that had pointed Gemini CLI at AGENTS.md needed no change.

2 more excerpts
  • Correction: Antigravity CLI reads AGENTS.md automatically by name with no settings key, so the earlier 're-check the flip' claim was broken and the source is used as a counter-example instead
  • Vendor migration guide asserting drop-in continuation; not independently tested

Gemini CLI to Antigravity CLI migration guide ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Birgitta Bockeler (martinfowler.com) Apr 2, 2026 practitioner

'In coding agents, part of the harness is already built in (e.g. via the system prompt, or the chosen code retrieval mechanism, or even a sophisticated orchestration system).'

The harness arrives partly pre-built - the inherited-defaults premise, from the discipline's highest-authority restatement.

4 more excerpts
  • The 2x2 that organizes the post: guides (feedforward) vs sensors (feedback), crossed with computational (deterministic, reliable) vs inferential (non-deterministic)
  • She files AGENTS.md and Skills as inferential feedforward - a legitimate quadrant member, not the harness's opposite
  • 'Building this outer harness is emerging as an ongoing engineering practice, not a one-time configuration'
  • Names cybernetics as the lineage via a Wikipedia link, with no specific control-theory originator

Harness engineering for coding agent users ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Cara Phillips, with Paul Chen, Andy Schumeister, Brad Abrams, Theo Chu (Anthropic) January 23, 2026 practitioner

Anthropic reports that multi-agent implementations typically use 3-10x more tokens than single-agent approaches for equivalent tasks. Three named conditions justify multiple agents: context pollution degrading performance, tasks that can run in parallel, and specialization that improves tool selection or task focus.

A closed list of three conditions is the only thing that justifies adding an agent; outside them, coordination costs typically exceed the benefits.

3 more excerpts
  • The 3-10x figure is unqualified internal testing with no published methodology, task list, dataset size, or model versions.
  • Coding-specific: dividing by type of work (one agent writes features, another writes tests, a third reviews code) creates constant coordination overhead; an agent handling a feature should also handle its tests.
  • The piece never mentions humans, approval, or review; verification is framed entirely as an agent-to-agent subagent pattern.

When to use multi-agent systems (and when not to) ↗·Cited in The Topology You Can Review Is the Topology You Can Run

Claude by Anthropic May 14, 2026 practitioner

The root file should be pointers and critical gotchas only; everything else drifts into noise.

The root tier's content rule comes from Anthropic itself: pointers and gotchas, not documentation.

2 more excerpts
  • Claude loads them additively as it moves through the codebase: root file for the big picture, subdirectory files for local conventions.
  • Skills solve this through progressive disclosure, offloading specialized workflows and domain knowledge that would otherwise compete for context space and loading them only when the task calls for it.

How Claude Code works in large codebases: best practices and where to start ↗·Cited in How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

Claude by Anthropic November 25, 2025 practitioner partial

Every conversation starts with this context already loaded, eliminating the need to explain basic project information repeatedly.

The always-loaded tier recurs every session — the recurring-cost premise. (Excerpt deliberately omits the page's 'system prompt' clause, refuted 0-3 against the docs.)

Using CLAUDE.md files: Customizing Claude Code for your codebase ↗·Cited in How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

Claude Code docs (code.claude.com) 2026 (undated on page) practitioner

'Claude Code reads CLAUDE.md, not AGENTS.md. If your repository already uses AGENTS.md for other coding agents, create a CLAUDE.md that imports it so both tools read the same instructions without duplicating them.' Memory files are treated as context, not enforced configuration; to block an action regardless of what Claude decides, use a PreToolUse hook.

The bridge for Claude Code is a one-line import, and the instruction file is advisory rather than enforcing.

4 more excerpts
  • CLAUDE.md content is delivered as a user message after the system prompt, not as part of the system prompt itself. Claude reads it and tries to follow it, but there's no guarantee of strict compliance, especially for vague or conflicting instructions.
  • CLAUDE.md and CLAUDE.local.md files in the directory hierarchy above the working directory are loaded in full at launch. Files in subdirectories load on demand when Claude reads files in those directories.
  • A symlink also works, but Windows symlinks need Administrator privileges or Developer Mode, so the @AGENTS.md import is the recommended route
  • Default /init reads Cursor and Copilot rules only; AGENTS.md, Devin, Windsurf and Cline rules are read only with CLAUDE_CODE_NEW_INIT=1

How Claude remembers your project ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness, CLAUDE.md Instruction Ceiling: Maintained Config, Not a README, How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

Claude Code docs (code.claude.com) practitioner

Subagent frontmatter carries permissionMode, hooks, mcpServers, maxTurns, memory and isolation (set to worktree to run in a temporary git worktree), and when the parent runs auto mode any permissionMode in the subagent frontmatter is ignored.

The subagent file format is markdown with YAML frontmatter, but the fields are Claude Code's own and do not map onto other tools.

4 more excerpts
  • Use one when a side task would flood your main conversation with search results, logs, or file contents you won't reference again: the subagent does that work in its own context and returns only the summary.
  • Definition precedence: managed settings, --agents CLI flag, .claude/agents/, ~/.claude/agents/, plugin agents; the definition closest to the working directory wins (v2.1.178+)
  • Subagent hooks support PreToolUse, PostToolUse and Stop, converted to SubagentStop at runtime
  • permissionMode: bypassPermissions is ignored when permissions.disableBypassPermissionsMode is set (v2.1.223+)

Create custom subagents ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness, When One Agent Stops Being Enough: The Isolation Gate

Claude Code docs (code.claude.com) undated (version gates reference up to v2.1.260) practitioner

'By default, if the sandbox cannot start because dependencies are missing or the platform is unsupported, Claude Code shows a warning and runs commands without sandboxing. To make this a hard failure instead, set sandbox.failIfUnavailable to true.' Native Windows is not supported.

The vendor sandbox fails open unless you configure it not to, so it cannot be the enforcement boundary.

4 more excerpts
  • 'The operating system enforces the sandbox boundary on the running process, so it holds regardless of what the model chose to run and even if an allowed command does more than its name suggests.'
  • Paths and domains from both sandbox settings and permission rules are merged into the final sandbox configuration, so the sandbox is interlocked with Claude Code's own permission grammar
  • Protected paths (.claude settings, skills, agents, commands, hooks, .mcp.json, credentials) cannot be exempted except by disabling filesystem isolation entirely
  • In a linked git worktree the sandbox allows writes to the main repository's shared .git directory except hooks/ and config

Configure the sandboxed Bash tool ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness, Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Claude Code docs (code.claude.com) undated (version notes run through v2.1.248) practitioner

'Hooks are user-defined shell commands, HTTP endpoints, MCP tool calls, LLM prompts, or subagents that execute automatically at specific points in Claude Code's lifecycle.' The reference lists 33 events, five handler types and seven configuration locations, and exit 2 means a blocking error that even a JSON permissionDecision of allow cannot override.

Hooks are a vendor-specific enforcement surface with their own event vocabulary, so they belong in the per-tool adapter pile.

4 more excerpts
  • 'For most hook events, only exit code 2 blocks the action. Claude Code treats exit code 1 as a non-blocking error and proceeds with the action, even though 1 is the conventional Unix failure code.'
  • Matcher semantics changed across point releases (v2.1.195, v2.1.214, v2.1.248), so hook configs are version-sensitive
  • The page says handlers run in the current directory with Claude Code's environment; it does not say 'full user permissions', treat that as inference
  • Cloud sessions on Claude Code on the web do not read local ~/.claude/settings.json; hooks there come from the repo and server-managed settings

Hooks reference ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness, Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Claude Code docs (code.claude.com) undated (version markers v2.1.172 to v2.1.239) practitioner

When using Amazon Bedrock, the /logout command is unavailable and the WebSearch tool is not available. Without pinning, model aliases such as sonnet and opus resolve to Claude Code's built-in default for Bedrock, which can lag the newest release, and Claude Code falls back to an earlier or lower-tier model at startup when the default is unavailable.

A second serving path for the same agent is a cheaper hedge than a second vendor, and it has a named feature cost.

3 more excerpts
  • Enabled with CLAUDE_CODE_USE_BEDROCK=1 plus AWS_REGION
  • A gateway that rewrites Content-Type on streaming responses forces a slower non-streaming path on every turn
  • The startup fallback is not persisted

Claude Code on Amazon Bedrock ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Claude Code docs (code.claude.com) practitioner

'The tradeoff is that the gateway becomes infrastructure your organization operates. Claude Code adds capabilities with each release, and a gateway that doesn't forward them breaks the corresponding features.' ANTHROPIC_BASE_URL is the variable that points Claude Code at the gateway.

Routing through a gateway buys provider switching and central credentials at the price of running infrastructure that must track the agent's release cadence.

3 more excerpts
  • Anthropic does not endorse, maintain or audit third-party gateways and does not support routing Claude Code to non-Claude models through any gateway
  • Provider switching without reconfiguring machines depends on the gateway exposing a single Anthropic-format endpoint
  • While a gateway credential variable or apiKeyHelper is active, a developer's claude.ai subscription is not used

Other LLM gateways ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Claude Code docs (code.claude.com) undated (version-gated content up to v2.1.257) practitioner

Six named modes (default, acceptEdits, plan, auto, dontAsk, bypassPermissions) plus a classifier. Setting auto or bypassPermissions in .claude/settings.json or .claude/settings.local.json does not take effect; if the classifier blocks an action 3 times in a row or 20 times total, auto mode pauses, and those thresholds are not configurable.

Claude Code's permission model is a product-specific state machine that cannot be expressed in another tool's config, and parts of it cannot be set from the repo at all.

4 more excerpts
  • Each action goes through a fixed decision order. The first matching step wins: allow/ask/deny rules resolve immediately, read-only actions and working-directory edits auto-approve second, everything else goes to the classifier third.
  • On Pro, Max and Team plans the built-in starting mode is auto (v2.1.228 or later); Enterprise, API-key, claude -p, Agent SDK and Bedrock sessions start in Manual
  • The classifier runs on Claude Sonnet 5 by default regardless of the /model selection
  • Under auto mode a subagent's permissionMode frontmatter is ignored

Choose a permission mode ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness, Where a Decision Model Belongs: Placing Jev in Your Stack

Claude Code docs (code.claude.com) practitioner

.mcp.json uses the standard top-level mcpServers shape and accepts streamable-http as an alias for http so configurations copied from server documentation work unchanged, but an entry with a url and no type is read as a stdio server and skipped with the error 'MCP server "<name>" has a "url" but no "type"'.

The mcpServers object is the portable element; scope precedence, approval gating and the type-field gotcha are Claude Code wiring.

3 more excerpts
  • Scope precedence: local, project, user, plugin-provided, claude.ai connectors; same-name servers are not merged
  • Project-scoped .mcp.json servers need approval before first use in interactive sessions; non-interactive contexts load them without prompting unless --strict-mcp-config (v2.1.246+) or disabledMcpjsonServers is used
  • anthropic/requiresUserInteraction in a tool's _meta forces an approval prompt on every call even under auto or bypassPermissions (v2.1.199+)

Connect Claude Code to tools via MCP ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Claude Code docs (code.claude.com) practitioner

A dev container 'is a convention rather than an enforcement boundary, because Claude Code does not require a container.' Under the built-in Bash sandbox, MCP servers and hooks are separate processes that run unconstrained on the host; the built-in Bash sandbox is the only approach Claude Code enforces itself.

Isolation that lives inside the vendor's product is a per-tool control surface; a container, VM, worktree or CI layer sits underneath every tool.

3 more excerpts
  • A sandboxed session that can write .mcp.json, .claude/commands or .claude/agents can persist hooks or MCP servers that run unsandboxed on the next launch
  • Isolation does not change what is sent to the model
  • Permission modes decide whether a call runs; isolation restricts what it can reach once it runs

Choose a sandbox environment ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Claude Docs (Anthropic) living docs (fetched Aug 3, 2026; inline version markers v2.1.198-v2.1.212) practitioner

'Running each Claude Code session in its own worktree means edits in one session never touch files in another.' The docs distinguish mechanisms explicitly: worktrees 'isolate file edits, while subagents and agent teams coordinate the work itself.'

File-ownership isolation is a shipped, first-class product mechanism - the enforcement layer exists today.

2 more excerpts
  • The page ships the enforcement mechanism but gives no guidance on how a human decides which files each session owns, and says nothing about what happens when two sessions need the same file - that absence is the gap the post addresses
  • Isolation is filesystem and branch level, not permission or prompt level: subagents take an 'isolation: worktree' frontmatter field, and Claude runs 'git worktree lock' while an agent is active

Run parallel sessions with worktrees ↗·Cited in Task Decomposition for AI Coding Agents: Draw the Graph First

Cloudflare Nov 18, 2025 practitioner partial

A change to a database system's permissions caused it to output multiple entries into a Bot Management feature file, which doubled in size and broke core traffic from 11:20 UTC, with recovery around 14:30 UTC and full normalization by 17:06 UTC.

Premise (b) of the shared-edge judgment: a shared edge provider failed the same day, in an overlapping window, from an internal change.

3 more excerpts
  • Names no AI vendor; the link to OpenAI's same-day incident is time-consistent but neither postmortem names the other
  • Explicitly not caused by an attack or malicious activity
  • No same-day Anthropic incident could be retrieved from the rolling status feed, so the shared-edge claim stays author's judgment

Cloudflare outage on November 18, 2025 ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Cursor (docs) undated (fetched Jul 27, 2026) practitioner partial

'Other exit codes - Hook failed, action proceeds (fail-open by default)'. The failClosed override ships with default false: 'When true, hook failures (crash, timeout, invalid JSON) block the action instead of allowing it through. Useful for security-critical hooks.'

A second vendor whose guardrail layer fails open by default, with the security switch shipped off.

3 more excerpts
  • Verification note: two independent passes disagreed on whether the page renders permission: "deny" or permission: 'deny'; that quote was dropped from the post rather than resolved
  • No raw-markdown endpoint (cursor.com/docs/hooks.md returns 404), so all quotes come from rendered HTML
  • beforeReadFile hook failures are logged and the read is allowed through

Hooks ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Cursor docs (cursor.com) practitioner

'Cursor supports AGENTS.md in the project root and subdirectories.' 'Unlike Project Rules, AGENTS.md is a plain markdown file without metadata or complex configurations.' Nested AGENTS.md files combine with parents, with more specific instructions taking precedence.

Cursor reads the shared file natively with no flip, while its .mdc rules and dashboard-managed Team Rules have no equivalent outside Cursor.

3 more excerpts
  • .mdc attachment logic (alwaysApply, globs, description-triggered) has no equivalent in a flat instructions file
  • Rules apply in the order Team Rules, Project Rules, User Rules; earlier sources win on conflict
  • Team and Enterprise plans can enforce rules from the Cursor dashboard so they cannot be disabled by members

Rules ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Cursor docs (cursor.com) practitioner

Twenty-one hook events (14 supported for cloud agents plus 7 not available to them) configured in .cursor/hooks.json, and 'Exit code 2 from command hooks blocks the action (equivalent to returning permission: "deny"). This matches Claude Code behavior for compatibility.'

Cursor is the only vendor doc that states its exit-2 contract matches Claude Code's, and even so the event list is its own.

2 more excerpts
  • Tab hooks and workspaceOpen are excluded from cloud agents because Tab completions are an IDE feature
  • Config paths: <project-root>/.cursor/hooks.json and ~/.cursor/hooks.json

Agent hooks ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Cursor docs (cursor.com) practitioner

Permissions are permissions.allow and permissions.deny arrays of rule strings such as Shell(rm), Read(.env*), Write(**/*.key) and Mcp(server:tool) in .cursor/cli.json or ~/.cursor/cli-config.json, and 'Deny rules take precedence over allow rules.'

Cursor's permission grammar is its own syntax in its own files, with no cross-tool equivalent.

2 more excerpts
  • Shell(commandBase) matches on the first token of the command line
  • Without an allowlist entry, each WebFetch prompts for approval

CLI permissions ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Cursor docs (cursor.com) practitioner partial

Cursor reads subagent files from .claude/agents/ (Claude compatibility) and .codex/agents/ (Codex compatibility) in addition to .cursor/agents/, with .cursor/ taking precedence on name conflicts; its own frontmatter fields readonly and is_background have no equivalent elsewhere.

Subagent non-portability is field-level only: the file location is interoperable, the fields (permissionMode, hooks, mcpServers on Claude; readonly, is_background on Cursor; TOML developer_instructions on Codex) are not.

2 more excerpts
  • Correction: the earlier claim that Cursor subagents do not map onto Claude Code or Codex was broken by this page; Cursor reads .claude/agents/ and .codex/agents/ explicitly
  • Frontmatter: name, description, model (default inherit), readonly (default false), is_background (default false)

Subagents ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Cursor docs (cursor.com) practitioner

Project MCP config lives in .cursor/mcp.json and global config in ~/.cursor/mcp.json, using the mcpServers shape, with Cursor-specific interpolation tokens such as ${env:NAME}, ${userHome}, ${workspaceFolder} and ${pathSeparator}.

The mcpServers object is the portable element; the file locations, interpolation and enterprise controls are Cursor wiring.

2 more excerpts
  • Enterprise Team Settings add command and URL entries, tool allowlists and network modes (Allow all, Allowlist, Deny all, No sandbox)
  • Transports: stdio, SSE and streamable HTTP

Model Context Protocol (MCP) ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Elasticsearch Labs (Someshwaran Mohankumar) January 16, 2026 practitioner

the most important memory work isn't 'store more,' it's 'curate better': Retrieve selectively, prune aggressively, summarize carefully

Reliability comes from curating context (selective retrieval, aggressive pruning), and tool-count bloat degrades even capable models.

2 more excerpts
  • once its context grew beyond a certain point (on the order of 100,000 tokens in an experiment), it began to fixate on repeating its past actions
  • failed a task when given 46 tools to consider but succeeded when given only 19 tools

Managing agentic memory with Elasticsearch ↗·Cited in Where Just-in-Time Context Retrieval Silently Breaks

Erik Schluntz, Barry Zhang (Anthropic) December 19, 2024 practitioner

Five named workflow patterns (Prompt chaining, Routing, Parallelization, Orchestrator-workers, Evaluator-optimizer), each with its own section and diagram. The only intermediate control point diagrammed as a node is a programmatic 'gate'; human checkpoints appear in a single line of prose.

The topology primitives predate later attempts to rename them, and no frontier-lab source elevates the human checkpoint to a named, diagrammed pattern the way it does the programmatic gate.

2 more excerpts
  • The subtractive bar, with 'only' italicized on the page: 'you should consider adding complexity only when it demonstrably improves outcomes.'
  • Anthropic gates 'multi-step agentic systems' while still endorsing multi-step workflows for well-defined tasks.

Building effective agents ↗·Cited in The Topology You Can Review Is the Topology You Can Run

Gemini CLI docs (geminicli.com) Jun 18, 2026 (last updated) practitioner

The context.fileName setting in settings.json ships an example list of 'AGENTS.md, CONTEXT.md, GEMINI.md', and large files can import others with the @file.md syntax using relative or absolute paths.

Gemini CLI is a one-key flip to read the shared AGENTS.md.

3 more excerpts
  • The page banner reads: 'Unpaid tier and Google One users: Gemini CLI was replaced by Antigravity CLI on June 18th, 2026', the same date as the last-updated stamp
  • Three-tier load order: global ~/.gemini/GEMINI.md, workspace ancestors, and just-in-time loading when a tool touches a directory
  • Gemini CLI itself is not durable for unpaid-tier readers; the post has to address that

Provide context with GEMINI.md files ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Gemini CLI docs (geminicli.com) practitioner

Eleven hook events across four groups: BeforeTool and AfterTool; BeforeAgent and AfterAgent; BeforeModel, BeforeToolSelection and AfterModel; SessionStart, SessionEnd, Notification and PreCompress. None shares a name with a Claude Code event.

Gemini CLI's hook vocabulary is disjoint from Claude Code's, so hooks are rewritten rather than copied.

3 more excerpts
  • The count of 11 is a manual tally across four section headings; there is no single master table
  • Blocking semantics differ per hook type: BeforeAgent aborts the turn and erases the prompt, AfterAgent rejects the response and triggers a retry
  • Configured in layered settings.json files (project .gemini/, user ~/.gemini/, system /etc/gemini-cli/)

Hooks reference ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Gemini CLI docs (geminicli.com) Apr 30, 2026 (last updated) practitioner

A priority-ordered rule engine: 'When a large language model wants to execute a tool, the policy engine evaluates all rules to find the highest-priority rule that matches the tool call', with allow, deny and ask_user decisions loaded from user, workspace (currently disabled) and admin TOML tiers, under an approval hierarchy of plan < default < autoEdit < yolo.

Gemini CLI's permission model is structurally unlike Codex's two flags and Claude Code's six modes, so there is no field-level translation.

3 more excerpts
  • ask_user is treated as deny in non-interactive mode
  • Workspace policies at $WORKSPACE_ROOT/.gemini/policies/*.toml are currently disabled; admin policies live in OS-specific system paths
  • Allow-for-all-future-sessions includes the current mode and all more permissive modes in the hierarchy

Policy engine ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Gemini CLI docs (geminicli.com) Jun 8, 2026 (last updated) practitioner partial

Custom agents are Markdown files with YAML frontmatter in .gemini/agents/ or ~/.gemini/agents/, with name and description required and fields including kind, tools, mcpServers, model, temperature, max_turns (default 30) and timeout_mins (default 10); the markdown body becomes the system prompt.

Same file shape as Claude Code's subagents, different directory and schema, so the definition is re-authored per tool.

3 more excerpts
  • model: inherit is not a documented value on this page
  • Subagents are treated as virtual tool names for policy matching, so access is governed by the TOML policy engine rather than by frontmatter
  • Tool wildcards: *, mcp_*, mcp_<server>_*

Subagents ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Gemini CLI docs (geminicli.com) Sep 2, 2026 (last updated) practitioner partial

'Gemini CLI uses the mcpServers configuration in your settings.json file to locate and connect to MCP servers.' The trust option 'bypasses all tool call confirmations for this server (default: false)' and the docs say to use it cautiously and only for servers you completely control.

The mcpServers block is the same shape across tools, while trust and allowlist merge rules are Gemini-specific.

3 more excerpts
  • The literal .gemini/settings.json path is not returned verbatim from this page; agents.md states it
  • excludeTools takes precedence over includeTools; when two sources both give an allowlist, only tools in both lists are enabled
  • Transports: stdio, SSE and streamable HTTP; fields include cwd, httpUrl, headers, timeout, targetAudience and targetServiceAccount

MCP servers with Gemini CLI ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Gemini CLI docs (geminicli.com) Mar 20, 2026 (last updated) practitioner

The vendor's own manual equivalent of its experimental --worktree flag is 'git worktree add ../project-feature-search -b feature-search' followed by cd and gemini; the flag stores worktrees under .gemini/worktrees/ and does not automatically delete the worktree or branch.

Worktree isolation is plain git, so it belongs in the layer that survives a vendor swap.

3 more excerpts
  • Gated behind an experimental opt-in: "experimental": { "worktrees": true } in settings.json
  • Cleanup is manual: git worktree remove .gemini/worktrees/<name> --force
  • The page carries the Gemini CLI to Antigravity CLI replacement banner

Git worktrees (experimental) ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

GitHub (anthropics/claude-code issue 6235) Opened Aug 21, 2025; closed Aug 17, 2026 practitioner

The year-long request (6,592 reactions, 389+ comments) was closed as completed by a comment pointing at the workaround: 'create a CLAUDE.md containing just @AGENTS.md (an import), or symlink CLAUDE.md to AGENTS.md.' A thread report from 2026-07-04 says a recent update blocked symlinked file updates with 'cannot write through symlinks'.

Native AGENTS.md support in Claude Code was declined in favor of the import, and the symlink route has broken in the field, which is why the post recommends import over symlink.

3 more excerpts
  • Closed with state_reason completed, not 'not planned' and not an inactivity auto-close
  • Maintainer status of the closing account (bcherny) was not independently verified beyond authoring the closing response
  • Community pushback the same night: the linked doc still requires a CLAUDE.md, a workaround rather than adoption

Feature Request: Support AGENTS.md ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

GitHub (anthropics/claude-code) June 25, 2025 practitioner

CLAUDE.md files in subdirectories are not being automatically loaded when accessing files in those directories, contrary to what the documentation states.

The lazy tier has unresolved reliability reports — documented design, verify on your surface (single macOS report, closed unresolved).

[BUG] CLAUDE.md files in subdirectories are not being automatically loaded (anthropics/claude-code#2571) ↗·Cited in How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

GitHub (anthropics/claude-code) opened Feb 11, 2026; closed not-planned Mar 22, 2026 practitioner

only the root-level CLAUDE.md is loaded at session start, and no subdirectory CLAUDE.md files are ever injected — even after multiple Read tool calls into those directories.

Lazy loading failed on the VS Code extension across three versions (2.1.39/2.1.45/2.1.49; CLI reportedly fine) — surface-specific reliability caveat.

[BUG] Subdirectory CLAUDE.md files not loaded on-demand when reading files via Read tool (anthropics/claude-code#24987) ↗·Cited in How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

GitHub (anthropics/claude-code) opened Mar 13, 2026; closed not-planned practitioner

The rules are clear in the CLAUDE.md and memory files — read them, I know them, and I still violated them. That's on me.

Anthropic closed 'ignored CLAUDE.md' as area:model, not-planned — structure is the practitioner's lever, not a forthcoming patch.

[BUG] Claude Code continually ignores CLAUDE.MD file (anthropics/claude-code#34197) ↗·Cited in How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

GitHub (anthropics/claude-code) opened Nov 16, 2025; closed not-planned practitioner partial

A modular CLAUDE.md structure with 6 referenced files (~2,100 lines total) consumes the same tokens as a monolithic file, providing organizational benefits only.

A user measured the @imports split delivering zero token savings, and Anthropic declined the lazy-imports request (author-self-reported measurement, not maintainer-confirmed).

1 more excerpt
  • 85-90% of loaded content is irrelevant to most conversations

[Feature Request] Lazy loading for @ file references in CLAUDE.md (anthropics/claude-code#11759) ↗·Cited in How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

GitHub Docs practitioner partial

'If multiple rulesets target the same branch or tag in a repository, the rules in each of these rulesets are aggregated' and 'the most restrictive version of the rule applies.' 'Anyone with read access to a repository can view its active rulesets.'

Enforcement on the host survives any agent swap because no vendor config can reach it, and it is auditable without admin access.

3 more excerpts
  • 'Evaluate mode' does not appear on this page; do not cite it to this URL
  • Up to 75 rulesets per repository and 75 organization-wide
  • Bypass can be granted to roles, teams or GitHub Apps when a ruleset is created; rulesets and branch protection rules enforce alongside each other

About rulesets ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

GitHub Docs practitioner

'By default, the restrictions of a branch protection rule don't apply to people with admin permissions to the repository or custom roles with the bypass branch protections permission' until 'Do not allow bypassing the above settings' is enabled; a required status check pinned to an app blocks merging 'if the status is set by any other person or integration.'

Two settings decide whether the host gate actually holds: closing the admin bypass and pinning the required check to its app.

3 more excerpts
  • Covers legacy branch protection rules; rulesets have their own bypass-actor model
  • Only one branch protection rule applies at a time, a restriction that does not apply to rulesets
  • Required checks must have a successful, skipped or neutral status before changes land

About protected branches ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Google (Gemini CLI docs) undated, main branch (fetched Jul 27, 2026) practitioner

'For global rules (those without an argsPattern), tools that are denied are completely excluded from the model's memory.' The model never sees the tool as an option.

A deny decision is not a request to the model - it is a change to what exists. Deterministic policy evaluation as reviewable code.

4 more excerpts
  • SCOPE: the memory-exclusion claim applies only to global rules without an argsPattern; the doc says nothing about argument-conditional deny rules
  • Precedence is arithmetic: final_priority = tier_base + (toml_priority / 1000), tier bases Default 1 through Admin 5 - tier always dominates because the fractional term can never reach 1
  • 'The first rule that matches determines the outcome' - first in priority order, not file order
  • In non-interactive mode ask_user is treated as deny

Policy engine ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Google Developers Blog May 19, 2026 (transition effective Jun 18, 2026) practitioner partial

'On June 18, 2026, Gemini CLI and Gemini Code Assist IDE extensions will stop serving requests for Google AI Pro and Ultra, as well as those using it free of charge using Gemini Code Assist for individuals.'

The one real vendor swap in the record: a vendor retired its own tool out from under individual users.

3 more excerpts
  • The blog names Google AI Pro and Ultra plus free individuals, not 'Google One'; that phrasing is the geminicli.com banner's
  • Organizations on Gemini Code Assist Standard or Enterprise licenses, or using Gemini Code Assist for GitHub through Google Cloud, keep unchanged access
  • Gemini Code Assist for GitHub also stopped new installations on June 18, 2026

An important update: Transitioning Gemini CLI to Antigravity CLI ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Guy (AWS Heroes) / DEV Community Posted Mar 18 (Edited Apr 8), 2026 practitioner

Tool descriptions are not documentation. They are the LLM's primary decision surface.

Tool descriptions are the agent's decision surface and must be audited like production code; nearly all of them carry quality defects.

3 more excerpts
  • 97.1% contain at least one quality issue
  • More than half (56%) have unclear purpose statements
  • augmented descriptions improved task success by 5.85 percentage points

MCP Tool Design: Why Your AI Agent Is Failing (And How to Fix It) ↗·Cited in Where Just-in-Time Context Retrieval Silently Breaks

Hacker News (bhaviav100, OP) 2026-03-29 (created_at 2026-03-29T00:22:20Z) practitioner

yes, compaction and smaller models help on cost per step. But my issue wasn't just inefficiency, it was agents retrying when they shouldn't. I needed visibility + limits per agent/task, and the ability to cut it off, not just optimize it.

Practitioners want per-agent/per-task limits and a hard cut-off, not just cost optimization — the wedge is attribute-and-enforce, not optimize.

4 more excerpts
  • My AGENTS.md is 845 lines and it only started getting good once it got that long" (Sammi), directly contested by "sweet spot is between 60 and 120 lines. With psuedo xml tags between sections" (typpilol)
  • Budget alerts are not a kill switch. Credits are not protection.
  • Claude often ignores CLAUDE.md / The more information you have in the file the more it gets ignored
  • "cost control is a policy problem - we certainly don't need to use opus 4.6 for a simple test refactor... we need a way to measure cost / performance for agents on individual repos, with individual types of tasks..." (author bisonbear, id 47563774)

Ask HN: How are you keeping AI coding agents from burning money? ↗·Cited in You Can't Cap What You Can't Attribute: Per-Task Cost, CLAUDE.md Instruction Ceiling: Maintained Config, Not a README

Harrison Chase (LangChain) June 16, 2025 practitioner

'The key insight is that read actions are inherently more parallelizable than write actions.' Parallelizing writes creates a dual problem: communicating context between agents and then merging their outputs coherently.

Whether an edge is safe to draw depends on whether the work crossing it reads or writes, which is why research-shaped multi-agent systems succeed where coding-shaped ones struggle.

2 more excerpts
  • Explicitly reconciles Cognition's 'Don't Build Multi-Agents' with Anthropic's research-system post as two correct answers for different work: 'Despite their opposing titles, I would argue they actually have a lot in common.'
  • Reproduces Anthropic's four-field subagent spec as an attributed block quote, so it is commentary on that material rather than independent corroboration of it.

How and when to build multi-agent systems ↗·Cited in The Topology You Can Review Is the Topology You Can Run

Ivan Kahl / Dometrain January 15, 2026 practitioner partial

You cannot craft the perfect CLAUDE.md file immediately. Instead, treat it as a living document.

CLAUDE.md is a living document refined over time, not a one-shot artifact.

2 more excerpts
  • Claude Code agents have a context window, and the CLAUDE.md file gets added to the agent's context. Any unnecessary instructions and wordy sentences will consume more of that context.
  • Always review the CLAUDE.md file and correct any assumptions or missing details related to project architecture.

Creating the Perfect CLAUDE.md for Claude Code ↗·Cited in CLAUDE.md Instruction Ceiling: Maintained Config, Not a README

Jude Gao (Vercel) Jan 27, 2026 practitioner partial

In 56% of eval cases the skill was never invoked. A compressed 8KB docs index embedded directly in AGENTS.md achieved a 100% pass rate while skills maxed out at 79%; skills without explicit instructions matched the no-docs baseline at 53%.

Whether an agent uses a portable artifact depends on the harness deciding to load it, which is a per-tool behavior rather than a property of the file.

2 more excerpts
  • Vendor eval on Next.js 16 API tasks; no inference-cost delta is reported on the page
  • The docs injection started at about 40KB and was compressed to 8KB, an 80% reduction

AGENTS.md outperforms skills in our agent evals ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Kyle (HumanLayer) November 25, 2025 practitioner

Frontier thinking LLMs can follow ~ 150-200 instructions with reasonable consistency.

There is a practical instruction ceiling — even frontier models only follow roughly 150-200 instructions consistently — so every line in CLAUDE.md competes for a finite budget.

4 more excerpts
  • At HumanLayer, our root CLAUDE.md file is less than sixty lines.
  • Claude Code's system prompt contains ~50 individual instructions
  • Smaller models get MUCH worse, MUCH more quickly
  • LLMs bias towards instructions that are on the peripheries of the prompt

Writing a good CLAUDE.md ↗·Cited in CLAUDE.md Instruction Ceiling: Maintained Config, Not a README, How to Structure CLAUDE.md: It's a Loading Policy, Not a Document

Lance Martin, Gabe Cemaj, Michael Cohen (Anthropic Engineering) Apr 8, 2026 practitioner

'Our only window in was the WebSocket event stream, but that couldn't tell us where failures arose, which meant that a bug in the harness, a packet drop in the event stream, or a container going offline all presented the same.'

A first-party operator account that having a telemetry stream is not the same as being able to localize a failure.

4 more excerpts
  • getEvents(), allows the brain to interrogate context by selecting positional slices of the event stream
  • Cited for the BEFORE state only. A full-text search confirms the article never reports an observability after-state
  • It names decoupling as the fix and scopes its durable event log to crash recovery and model-side context interrogation
  • Its quantified improvements (p50 TTFT down roughly 60%, p95 over 90%) are latency, not observability

Scaling Managed Agents: Decoupling the brain from the hands ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back, Loop Engineering Breaks Your Single-Shot Context Playbook

LangChain December 14, 2024 practitioner

The pause is implemented on the graph's own persistence layer: 'Every step of the graph, it reads from and then writes to a checkpoint of that graph state', which is what allows the run to 'pause execution of the graph half way through, and then resume after some time'.

The human checkpoint is durable because it rides the graph's existing checkpoint mechanism rather than an external callback, which is what makes it structurally different from a manual review step in a runbook.

1 more excerpt
  • Named human-in-the-loop patterns: Approve or Reject, Review and Edit State, Review Tool Calls, and multi-turn conversation in a multi-agent setup.

Making it easier to build human-in-the-loop agents with interrupt ↗·Cited in The Topology You Can Review Is the Topology You Can Run

LangChain (LangGraph OSS Python documentation) practitioner

'The interrupt function pauses graph execution and returns a value to the caller. When you call interrupt within a node, LangGraph saves the current graph state and waits for you to resume execution with input.' Named patterns include approval workflows, review and edit state, interrupts in tools, and validating human input.

A human checkpoint can be a first-class graph node with durable state rather than a process stage described in prose - the mechanism ships today in a production framework.

3 more excerpts
  • Durability is explicit: state is saved by the checkpointer, which in production should be database-backed.
  • Idempotency gotcha: 'The node restarts from the beginning... so any code before the interrupt runs again.'
  • The documentation contains no notion of review capacity, cost, or how many checkpoints is too many.

Interrupts ↗·Cited in The Topology You Can Review Is the Topology You Can Run

Lilian Weng (Lil'Log) Jul 4, 2026 practitioner

'The evaluator and permission control should likely sit outside the loop that evolves harness, with held-out tests, trace audits, and human review at decision points that matter.'

Enforcement belongs outside the loop it governs - and, because harnesses evolve, the audit is recurring rather than one-time.

4 more excerpts
  • SCOPE: scoped specifically to self-modifying harnesses, not agent harnesses generally; the 'should likely' hedge is the author's own
  • 'A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results'
  • The words deterministic and probabilistic appear nowhere in the post - do not attribute that framing to her
  • 'The core interface of mainstream coding agents has become stabilized across Claude Code, Codex, OpenCode, and Cursor-style agents'

Harness Engineering for Self-Improvement ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Medium (Paolo Perrone / Data Science Collective) April 12, 2026 practitioner

This isn't a hallucination. The retrieval worked perfectly. It just retrieved garbage.

Bad retrieval is a distinct silent failure mode from hallucination and has no built-in flag, so leaders must add one.

3 more excerpts
  • Silent retrieval failure. There's no mechanism to flag 'this retrieval returned low-confidence or low-credibility results.'
  • For 28 minutes, 55% of API requests to the platform failed
  • An agent with 85% accuracy per step only completes a 10-step workflow successfully 20% of the time

Why AI Agents Keep Failing in Production ↗·Cited in Where Just-in-Time Context Retrieval Silently Breaks

Molisha Shah (Augment Code) March 16, 2026 (updated June 18, 2026) practitioner partial

Six patterns for parallel coding agents, including git-worktree isolation (each agent gets its own working directory and index while sharing one .git object database) and sequential merges. 'Git detects textual conflicts, not semantic ones. Authoritative guidance still converges on mandatory human review for logic-level contradictions.'

Mechanical isolation between parallel coding agents solves textual collision but not semantic conflict, which is what leaves a human judgment step irreducible at the merge boundary.

2 more excerpts
  • Vendor guide authored by a go-to-market staff member rather than an engineer; the technical core is real but a '40% reduction in hallucinations' claim about the vendor's own product is unverified marketing and is not cited.
  • Humans are given a tier in a review pipeline ('Reserve humans for semantic correctness and architecture'), not a position in a graph.

How to Run a Multi-Agent Coding Workspace (2026) ↗·Cited in The Topology You Can Review Is the Topology You Can Run

OpenAI undated, on or before Feb 11, 2026 (Wayback-bounded) practitioner partial

'Agents are most effective in environments with strict boundaries and predictable structure, so we built the application around a rigid architectural model.'

The canonical build-a-harness text - and the differentiation foil: written from an empty repository, with zero coverage of inherited vendor defaults.

4 more excerpts
  • Direct fetch returns HTTP 403 and web.archive.org is blocked to the tool; content read via a text-extraction proxy
  • Page carries no byline and no publication date - attribute to OpenAI, not to an individual
  • Its 'boundaries' are architectural layer-dependency lint rules inside the application codebase, not harness control surfaces
  • Posture cuts against gates: 'The repository operates with minimal blocking merge gates'

Harness engineering: leveraging Codex in an agent-first world ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

OpenAI practitioner

OpenAI's stated default: 'Start with one agent whenever you can. Add specialists only when they materially improve capability isolation, policy isolation, prompt clarity, or trace legibility. Splitting too early creates more prompts, more traces, and more approval surfaces without necessarily making the workflow better.'

Both major labs publish a single-agent default with an explicit justification bar, and OpenAI counts human approval surface as one of the costs that premature decomposition imposes.

3 more excerpts
  • 'Approval surfaces' is used once and never defined, on this page or the parent guide.
  • Neither page uses 'review capacity' or 'reviewer bandwidth', and neither frames human review as a binding constraint on agent count - that framing belongs to the post, not to OpenAI.
  • 'Trace legibility' is listed as a reason to split, which cuts against reading this as a blanket anti-splitting stance.

Orchestration and handoffs ↗·Cited in The Topology You Can Review Is the Topology You Can Run

OpenAI (Codex docs) undated (fetched Jul 27, 2026) practitioner

'Some specialized tool paths can opt out of the default hook path. Treat tool hooks as a useful guardrail, not a complete enforcement boundary.'

A vendor stating in its own documentation that the layer most teams build policy on is not an enforcement boundary.

4 more excerpts
  • Codex hooks fire at 12 lifecycle points (PreToolUse, PermissionRequest, PostToolUse, PreCompact, PostCompact, UserPromptSubmit, SubagentStop, Stop, Interrupt, SessionStart, SubagentStart, SessionEnd), are configured in .codex/hooks.json or .codex/config.toml, and 'you can also use exit code 2 and write the blocking reason to stderr.'
  • The named hole: 'Hosted tools, such as WebSearch... don't use the local function-tool hook path'
  • Codex also fails open: a PreToolUse hook returning unsupported fields is marked failed and 'continues the tool call'
  • 'Multiple matching command hooks for the same event are launched concurrently, so one hook can't prevent another matching hook from starting'

Hooks ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews, AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

OpenAI (status.openai.com) Jun 9 to Jun 10, 2025 practitioner partial

ChatGPT users experienced elevated error rates reaching about 35% at peak while API error rates peaked near 25%, and API availability dropped to 75%; the absence of break-glass tooling to rapidly restore network connectivity on affected nodes extended the overall recovery timeline.

The second vendor has its own outage character: a routine host-OS update on GPU nodes fanned out into a roughly 15-hour cross-product outage.

3 more excerpts
  • The write-up does not mention Codex anywhere; only ChatGPT and the API are named, so the earlier 'Codex listed as affected' claim was dropped
  • Root cause was a systemd-networkd restart conflicting with a production networking agent, which removed all routes from affected nodes
  • Remediation disabled automatic daily updates on GPU systems

Elevated error rates ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

OpenAI (status.openai.com) Nov 18, 2025 practitioner

'We have confirmed that the incident is caused by an issue with one of our third-party service providers.' The outage affected APIs and ChatGPT from 12:12 PM to 3:18 PM, roughly 3 hours 6 minutes.

Premise (a) of the shared-edge judgment: OpenAI attributed a multi-hour outage to an unnamed upstream provider.

3 more excerpts
  • No provider is named; do not assert Cloudflare from this source
  • The promised root cause analysis was not located
  • Over an hour passed before the cause category was named

Access issues affecting OpenAI websites ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

OpenAI Codex docs (learn.chatgpt.com) practitioner

The import maps instruction files to AGENTS.md, settings.json to config.toml, slash commands to Skills, subagents to Codex subagents, MCP configuration to Codex MCP configuration and hooks to Codex hooks, then lists tool restrictions and permissions, MCP auth and headers, hook behavior, plugins and argument-bearing prompts as items to review manually.

The vendor's own migration tool draws the line between what carries over and what has to be rebuilt per tool.

3 more excerpts
  • Supported sources: Claude Code, Claude Cowork and Cursor for the desktop app; Claude Code and Cursor for Codex CLI
  • The page describes importing into Codex only, never exporting; use 'items to review', not 'silently converted'
  • Chats import at most 50 from the last 30 days, which is state rather than configuration

Import from another agent ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

OpenAI Codex docs (learn.chatgpt.com) practitioner

Codex checks each directory in this order: AGENTS.override.md, AGENTS.md, TEAM_GUIDE.md, .agents.md (fallbacks set via project_doc_fallback_filenames), uses only the first non-empty file per level, and stops adding files once the combined size reaches project_doc_max_bytes, default 32 KiB.

Codex reads AGENTS.md natively and rebuilds the instruction chain root-down each run, so no flip is required.

2 more excerpts
  • Files closer to the current directory override earlier guidance because they appear later in the chain
  • Spot-check the literal fallback example array before quoting

Custom instructions with AGENTS.md ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

OpenAI Codex docs (learn.chatgpt.com) practitioner

Permissions are a cross-product of approval_policy (untrusted, on-request, never, or granular) and sandbox_mode (read-only, workspace-write, danger-full-access), where danger-full-access disables sandboxing entirely; the combined --dangerously-bypass-approvals-and-sandbox flag is aliased --yolo.

Codex's two-axis permission model has no field-level equivalent to Claude Code's six named modes, so permissions do not port.

2 more excerpts
  • Codex loads project-scoped .codex/ config layers only when the project is trusted and may start read-only until the working directory is trusted
  • The six named sandbox-and-approval combinations come from the companion agent-approvals-security page

Advanced configuration ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

OpenAI Codex docs (learn.chatgpt.com) practitioner

Every standalone custom agent file must define name, description and developer_instructions, lives as TOML under .codex/agents/ or ~/.codex/agents/, and the docs say 'the format may evolve as authoring and sharing mature.'

The Codex subagent format is TOML with its own required keys, not the markdown-plus-frontmatter shape Claude Code, Gemini CLI and Cursor share.

3 more excerpts
  • Codex reapplies the parent turn's live runtime overrides (such as /permissions changes or --yolo) when spawning a child, even if the agent file sets different defaults
  • Settings resolve from an explicit spawn value, then the [agents] default, then the parent's value
  • Other config.toml keys such as model, sandbox_mode, mcp_servers and skills.config can be included in an agent file

Subagents ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

OpenAI Codex docs (learn.chatgpt.com) practitioner

MCP servers are configured as [mcp_servers.<server-name>] TOML tables in ~/.codex/config.toml or a project-scoped .codex/config.toml (trusted projects only), and 'the ChatGPT desktop app, Codex CLI, and IDE extension share this configuration.'

Codex shares MCP config across OpenAI's own surfaces only, in TOML rather than the JSON mcpServers shape the other three tools use.

3 more excerpts
  • Per-server default_tools_approval_mode accepts auto, prompt, writes and approve, plus per-tool approval_mode
  • Codex validates any returned iss before exchanging an OAuth authorization code; a mismatch always rejects
  • Codex reads the MCP instructions field returned at initialization as server-wide guidance

Model Context Protocol ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

OpenAI Codex docs (learn.chatgpt.com) practitioner

Codex creates worktrees in $CODEX_HOME/worktrees in a detached HEAD state, keeps the most recent 15 Codex-managed worktrees by default, and copies ignored files only if they match .worktreeinclude (AGENTS.override.md is copied automatically).

Vendor-managed worktrees are a convenience over plain git worktree that is scoped to one product's parallel-chat feature.

2 more excerpts
  • The mechanism is scoped to the Codex desktop app running parallel chats
  • AGENTS.override.md is a Codex-specific override file distinct from the AGENTS.md convention

Worktrees ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

OpenAI Codex docs (learn.chatgpt.com) practitioner

'Pair AGENTS.md with infrastructure that enforces those rules: pre-commit hooks, linters, and type checkers catch issues before you see them.' Files closer to the working directory take precedence, and nested AGENTS.md supersede parents while the repo overrides ~/.codex/AGENTS.md.

The vendor's own advice is that the instruction file is not the enforcement layer; enforcement lives in tooling no agent can talk past.

2 more excerpts
  • Skills live at ~/.agents/skills (global) and .agents/skills (repo) as SKILL.md with optional scripts, references and assets
  • agents/openai.yaml (MCP dependencies for skills) and Plugins are explicitly OpenAI/Codex-specific

Customization ↗·Cited in AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness

Simon Willison Sept 30, 2025 practitioner

A critical new skill to develop is designing agentic loops.

Designing the loop an agent runs in is a distinct skill that predates the 'loop engineering' label by nine months.

1 more excerpt
  • My preferred definition of an LLM agent is something that runs tools in a loop to achieve a goal.

Designing agentic loops ↗·Cited in Loop Engineering Breaks Your Single-Shot Context Playbook

Simon Willison Mar 16, 2026 practitioner

'A coding agent is a piece of software that acts as a harness for an LLM, extending that LLM with additional capabilities that are powered by invisible prompts and implemented as callable tools.'

The harness's prompts are invisible to the user by construction - the closest practitioner framing to the unreviewed-defaults claim.

1 more excerpt
  • 'A tool is a function that the agent harness makes available to the LLM'

How coding agents work ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Simon Willison Jun 16, 2025 practitioner

'you can try telling it not to in your own prompt, but how confident can you be that your protection will work every time?'

A system-prompt instruction is not a security control - the cleanest one-line statement of the post's premise.

1 more excerpt
  • On vendor guardrails claiming '95% of attacks': 'in web application security 95% is very much a failing grade' - note the 95% is Willison characterizing vendor marketing, not his own measurement

The lethal trifecta for AI agents: private data, untrusted content, and external communication ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Simon Willison's Weblog 27th June 2025 practitioner

context engineering is the delicate art and science of filling the context window with just the right information for the next step...task descriptions and explanations, few shot examples, RAG, related (possibly multimodal) data, tools, state and history

Context engineering, not prompt engineering, is the real discipline: filling the window with the right information environment for the next step.

1 more excerpt
  • the art of providing all the context for the task to be plausibly solvable by the LLM

Context engineering ↗·Cited in Where Just-in-Time Context Retrieval Silently Breaks

Towards Data Science (Mostafa Ibrahim) March 20, 2026 practitioner

The agent optimises locally. At each step, it asks, 'Do I have enough?' and when the answer is uncertain, it defaults to 'get more'. Without hard stopping rules, the default spirals.

Without a hard stop rule, an agent's local 'get more' default turns retrieval into an unbounded budget fire; capping cycles and abstaining is the control.

3 more excerpts
  • Three cap retrieval cycles. After three failed passes, return a best-effort answer with a confidence disclaimer.
  • agents making 200 LLM calls in 10 minutes, burning $50–$200 before anyone noticed
  • costs spike 1,700% during a provider outage as retry logic spiralled out of control

Agentic RAG Failure Modes: Retrieval Thrash, Tool Storms, and Context Bloat (and How to Spot Them Early) ↗·Cited in Where Just-in-Time Context Retrieval Silently Breaks

TrueFoundry June 18, 2026 practitioner

The honest column in the ledger: JIT buys these at the price of retrieval latency on the steps that load (usually trivial next to a model call, but nonzero), a new failure mode (an unresolvable reference must surface as an honest error, not a hallucinated payload), and a dependency on description quality — the agent loads from the catalog's one-liners, so a bad stub hides a good payload.

JIT context introduces two specific liabilities leaders must design for: unresolvable references must fail loud as honest errors, and retrieval quality is capped by the quality of catalog descriptions.

3 more excerpts
  • in a loop, the window is re-sent every step, so a preloaded handbook isn't one payment but thirty
  • a preloaded copy is a snapshot that ages as the run proceeds, while a reference resolves to the current state of the file, the ticket, the database at the moment of use
  • long-context research and practitioner experience agree that models degrade as windows fill with low-relevance text

JIT Context: Why the Best Agents Load Late and Load Little ↗·Cited in Where Just-in-Time Context Retrieval Silently Breaks

Truong Phung (DEV Community) Jun 25, 2026 practitioner

The model forgets everything between runs, so state must live on disk, not in the context window. The agent forgets; the repo doesn't.

Cross-iteration state must live on disk, and every unattended loop needs three hard stops: max iterations, no-progress detection, and a spend budget.

1 more excerpt
  • Bake all three into every loop: 1. Max iteration count 2. No-progress detection 3. A token or dollar budget

The Agentic Loop / Loop Engineering: A Practical Field Guide ↗·Cited in Loop Engineering Breaks Your Single-Shot Context Playbook

Vivek Trivedy Mar 10, 2026 practitioner

'A harness is every piece of code, configuration, and execution logic that isn't the model itself. A raw model is not an agent. But it becomes one when a harness gives it things like state, tool execution, feedback loops, and enforceable constraints.'

The canonical definition, from the coinage. Cite Trivedy rather than downstream restatements.

1 more excerpt
  • The component enumeration explicitly places System Prompts inside the harness alongside tools, sandbox, orchestration logic, and hooks - so 'harness vs CLAUDE.md' is not the boundary the field draws

The Anatomy of an Agent Harness ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Zhu Liang August 5, 2025 practitioner

negative instructions can be unreliable as user prompts

Negative 'don't do that' rules are unreliable in a user message like CLAUDE.md, so positive, runnable framing is preferable — reserving DO-NOT for hard safety boundaries.

3 more excerpts
  • Reddit user reported Claude Code created duplicate files despite explicit 'NEVER create duplicate files' rule
  • Gemini models have 'hit-or-miss' performance with negative commands
  • They are effective at preventing unethical or harmful behavior, especially when used in system prompts

The Pink Elephant Problem: Why 'Don't Do That' Fails with LLMs ↗·Cited in CLAUDE.md Instruction Ceiling: Maintained Config, Not a README

Evals & Verification

44 sources

Knowing an agent’s output is actually correct, beyond a green build.

Albayaydh, Zhao, Flechais Jul 7, 2026 data

A synthesis of 27 benchmark, taxonomy, and audit papers spanning 19 benchmarks finds that additional scaffolding does not consistently improve reliability.

The honesty brake on 'more harness is better' - and the reason this post's claim is bounded to blast radius rather than quality.

3 more excerpts
  • Failures compound nonlinearly with task length
  • Strong performance on individual sub-tasks does not reliably translate into end-to-end success
  • Secondary synthesis - every number in it is someone else's measurement

Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Aleithan, Xue, Mohajer, Nnorom, Uddin, Wang October 9, 2024 data

32.67% of successful SWE-bench patches involved solution leakage (the fix present in the issue report or comments) and 31.08% passed on weak tests; filtering both drops SWE-Agent+GPT-4's resolution rate from 12.47% to 3.97%.

A third of measured SWE-bench success was answer leakage, a concrete mechanism by which leaderboard scores inflate without capability.

1 more excerpt
  • Over 94% of benchmark issues predate LLM knowledge cutoff dates

SWE-Bench+: Enhanced Coding Benchmark for LLMs ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Andrew Gelman, Eric Loken (Columbia University / Penn State) November 14, 2013 data

Researcher degrees of freedom can produce a multiple-comparisons problem even where researchers perform only a single analysis on their data, because the analysis actually run is one of many that would have been equally reasonable had the data come out differently.

A team reading its own agent-PR dashboard once, in good faith, with a hypothesis stated in advance, is already exposed to the fragility — no p-hacking or repeated analysis is required for the result to be contingent on arbitrary choices.

2 more excerpts
  • 'The researcher degrees of freedom do not feel like degrees of freedom because, conditional on the data, each choice appears to be deterministic'
  • Cited from the 2013 unpublished manuscript; the 2014 American Scientist version ('The Statistical Crisis in Science') was not reachable for verification and is not the version quoted

The garden of forking paths: Why multiple comparisons can be a problem, even when there is no "fishing expedition" or "p-hacking" and the research hypothesis was posited ahead of time ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Anthropic Engineering January 9, 2026 data

Teams delay building evals thinking they need hundreds of tasks; in reality 20-50 simple tasks drawn from real failures is a great start, structured by task/trial/outcome vocabulary.

A working internal agent eval suite is a 20-50 task project, not an infrastructure program - removing the main excuse for deciding from public leaderboards instead.

4 more excerpts
  • Opus 4.5 initially scored 42% on CORE-Bench; after fixing grading bugs and using a less constrained scaffold, the same model's score jumped to 95%.
  • So as not to unnecessarily punish creativity, it's often better to grade what the agent produced, not the path it took.
  • 'With frontier models, a 0% pass rate across many trials (i.e 0% pass@100) is most often a signal of a broken task, not an incapable agent'
  • The initial 42% observation is Anthropic citing an external report; the diagnosis and the 95% re-run are Anthropic's own

Demystifying evals for AI agents ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard, Audit Your Agent Harness: The Deterministic Layer Nobody Reviews, How to Verify AI Coding Agent Output: A Reviewer's Framework

Anthropic Engineering February 5, 2026 data

In internal experiments spanning six compute-resource configurations on GKE with model, harness, and task set held fixed, the gap between the most- and least-resourced setups on Terminal-Bench 2.0 was 6 percentage points (p < 0.01); infrastructure error rates fell from 5.8% under strict 1x enforcement to 2.1% at 3x headroom and 0.5% uncapped.

Infrastructure configuration alone produces score differences exceeding the few-point margins that separate top leaderboard entries, so cross-infrastructure leaderboard comparisons are not decision-grade evidence for a model swap.

4 more excerpts
  • Infrastructure configuration can swing agentic coding benchmarks by several percentage points - sometimes more than the leaderboard gap between top models. The gap between most- and least-resourced setups on Terminal-Bench 2.0 was 6 percentage points (p < 0.01).
  • A 2-point lead on a leaderboard might reflect a genuine capability difference, or it might reflect that one eval ran on beefier hardware, or even at a luckier time of day, or both.
  • Top leaderboard spots are often separated by just a few percentage points, per the post's own framing
  • Resource headroom is an eval-infrastructure design requirement: error rate falls an order of magnitude from strict to uncapped provisioning

Quantifying infrastructure noise in agentic coding evals ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard, Audit Your Agent Harness: The Deterministic Layer Nobody Reviews, How to Verify AI Coding Agent Output: A Reviewer's Framework

arXiv (Ahmed E. Hassan, Hao Li, Dayi Lin, Bram Adams, Tse-Hsun Chen, Yutaro Kashiwa, Dong Qiu) 2025 data

Their hyper-productivity is revealing a significant 'speed vs. trust' gap. Recent, deeper examinations of agent-generated code and agent-driven PRs reveal that a large percentage of agent efforts fail to meet the quality bar of being truly 'merge-ready,' often containing subtle regressions, superficial fixes, or a general lack of engineering hygiene.

Agent hyper-productivity creates a speed-vs-trust gap where most agent PRs aren't merge-ready, overwhelming review capacity.

3 more excerpts
  • 29.6% of 'plausible' fixes introduced behavioral regressions or were incorrect upon rigorous retesting
  • True solve rates for GPT-4 patches dropped from 12.47% to 3.97% after detailed manual audits
  • Over 68% of agent-generated pull requests reportedly face long delays or remain unreviewed, creating an urgent need for scalable review automation.

Agentic Software Engineering: Foundational Pillars and a Research Roadmap ↗·Cited in How to Verify AI Coding Agent Output: A Reviewer's Framework

arXiv (Jingzhi Gong, Giovanni Pinna, Yixin Bian, Jie M. Zhang) Submitted January 8, 2026; revised January 26, 2026 data

descriptions claim unimplemented changes" was the most common issue (45.4%); high-MCI PRs had 51.7% lower acceptance rates (28.3% vs. 80.0%)

The most common defect in agent-authored PRs is a description claiming changes the code never made, and those PRs get accepted far less.

3 more excerpts
  • 23,247 agentic PRs analyzed across five agents
  • High-MCI PRs took 3.5 times longer to merge (55.8 vs. 16.0 hours)
  • 406 PRs (1.7%) exhibited high PR-MCI

Analyzing Message-Code Inconsistency in AI Coding Agent-Authored Pull Requests ↗·Cited in How to Verify AI Coding Agent Output: A Reviewer's Framework

arXiv (Sabrina Haque, Sarvesh Ingale, Christoph Csallner) Submitted January 7-8, 2026 data partial

Across agents, test-containing PRs are more common over time and tend to be larger and take longer to complete, while merge rates remain largely similar.

Whether an agent PR includes tests varies and doesn't correlate with merge outcomes, so test presence is a signal to read, not proof of quality.

2 more excerpts
  • We observe variation across agents in both test adoption and the balance between test and production code within test PRs
  • Testing is a critical practice for ensuring software correctness and long-term maintainability

Do Autonomous Agents Contribute Test Code? A Study of Tests in Agentic Pull Requests ↗·Cited in How to Verify AI Coding Agent Output: A Reviewer's Framework

arXiv (Yue Liu, Ratnadira Widyasari, Yanjie Zhao, Ivana Clairine Irsan, Junkai Chen, David Lo) Submitted 30 March 2026 (v2 revised 26 April 2026) data partial

22.7% of tracked AI-introduced issues still survive at the latest version of the repository. These findings show that AI-generated code can introduce long-term maintenance costs into real software projects.

Over a fifth of AI-introduced issues survive at HEAD, so AI code accrues durable technical debt at scale unless verification catches it.

3 more excerpts
  • 302.6k verified AI-authored commits from 6,299 GitHub repositories
  • more than 15% of commits from every AI coding assistant introduce at least one issue
  • code smells are by far the most common type" / "89.3% of all issues

Debt Behind the AI Boom: A Large-Scale Empirical Study of AI-Generated Code in the Wild ↗·Cited in How to Verify AI Coding Agent Output: A Reviewer's Framework

Berkeley RDI (Hao Wang, Qiuyang Mang, Alvin Cheung, Koushik Sen, Dawn Song) April 2026 data

We built an automated scanning agent that systematically audited eight among the most prominent AI agent benchmarks — SWE-bench, WebArena, OSWorld, GAIA, Terminal-Bench, FieldWorkArena, and CAR-bench — and discovered that every single one can be exploited to achieve near-perfect scores without solving a single task.

Every major agent benchmark can be gamed to near-perfect scores without solving anything, so self-reported benchmark performance is structurally untrustworthy.

4 more excerpts
  • Every one of eight major AI agent benchmarks audited can be exploited to near-perfect scores without solving a single task - including 100% on SWE-bench Verified via a 10-line conftest.py that hooks pytest and rewrites every test result to passed.
  • A conftest.py file with 10 lines of Python 'resolves' every instance on SWE-bench Verified.
  • SWE-bench Verified (500 tasks) — 100% score via pytest hooks
  • Benchmark scores are actively being gamed, inflated, or rendered meaningless, not in theory, but in practice.

How We Broke Top AI Agent Benchmarks: And What Comes Next ↗·Cited in How to Verify AI Coding Agent Output: A Reviewer's Framework, Stop Picking Your Coding-Agent Model Off a Leaderboard

Bitya Neuhof, Yuval Benjamini June 28, 2026 data

On MMLU (57 subjects), three distinct models can be ranked as the fourth from the top - statistically consistent with the same rank position - with uncertainty driven more by between-subject variability than prompt variants.

Leaderboard ranks published as single values hide that several models are often statistically tied for the same position.

Quantifying Ranking Uncertainty in LLM Benchmarks ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Carlo A. Furia, Richard Torkar January 28, 2025 (latest revision April 1, 2026) data partial

In a study of 105 projects (45 Java, 60 Python) and 952 developers, omitting programmer skill inflated the measured programming-language effect on code quality roughly fourfold, from an adjusted -0.012 to a confounded -0.052. Sensitivity analysis showed a scaled-mean difference of -0.062 would be sufficient to flip the sign of the measured effect.

An aggregate software-engineering trend can be substantially distorted — and, at plausible confounding strengths, sign-flipped — by omitting a variable that is a real determinant of the outcome, which is what a dashboard that does not condition on team properties is doing.

3 more excerpts
  • 'Additional data about X and Y (i.e., sampling more datapoints) is not going to help; in fact, it may just entrench our reliance on the biased estimate by reducing its variance and giving the false impression of reliability'
  • The sign flip is a sensitivity-analysis threshold showing how little confounding would suffice, not an observed reversal in the data; neither case study concerns code review
  • Preprint with no journal-ref; Simpson's paradox is not discussed anywhere in the paper

Mitigating Omitted Variable Bias in Empirical Software Engineering ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Christopher Kelly, Angelica Chowdhury, Alexandra Campili, Bimpe Ayoola, Devin Barbour, Thomas Chen Dawson, Ze Shen Chin, Rokas Gipiškis May 3, 2026 (revised July 26, 2026; ICML 2026 Workshop on Technical AI Governance) data

Synthesizes five principles from established validity frameworks into 33 guidelines for AI evaluation trials, naming construct underrepresentation and construct-irrelevant variance as the failure modes that make AI-effect measurements uninterpretable.

Where construct validity is weak, a study can show that scores increased with AI without credibly claiming the underlying capability improved — which is why 'is AI good for code review' is not merely contested but malformed as a measurable question.

2 more excerpts
  • Guideline 18 is directly applicable to reading a vendor claim: 'Apply identical quality rubrics, performance benchmarks, and scoring criteria to human-only and human+AI outputs'
  • Software engineering appears as a donor methodology field and 'coding competence' as an example construct; the paper analyzes no code-review study

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Evan Miller (Anthropic) November 1, 2024 data

Clustered standard errors on public evals can be over 3X larger than naive standard errors, and detecting an absolute score difference of 0.03 at 80% power requires an eval of at least ~969 independent questions.

Most reported model-to-model benchmark gaps are narrower than honestly computed confidence intervals - the statistical foundation for treating small leaderboard gaps as ties.

2 more excerpts
  • The same pair of models can differ significantly on one benchmark (MATH) and not on others (HumanEval, MGSM) in the paper's worked example
  • The Llama 3 paper's reported confidence intervals are judged likely anti-conservative (too narrow)

Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Florian Brand, Jean-Stanislas Denain (Epoch AI) June 13, 2025 data

A good scaffold can increase SWE-bench Verified performance by up to 20%, so scores reflect the sophistication of the scaffold as much as the capability of the underlying model.

Scaffold quality is a confound baked into every SWE-bench Verified score - the leaderboard measures a model-plus-scaffold system.

1 more excerpt
  • Concrete scaffold example: SWE-Agent's 100-line file viewer, linter-integrated edit tool, and custom directory search

What skills does SWE-bench Verified evaluate? ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Guo, Liu, Zhang, Ma, Lou, Chen July 1, 2026 data

Across five runs, run-to-run standard deviations were 2.0 percentage points for SWE-Doctor (the most stable agent), 2.2 for mini-SWE-agent, and 3.4 for live-SWE-agent; SWE-Doctor's Pass@5 was 70.0% against All@5 of 40.0%.

Even the most stable SWE-bench-family agents swing multiple percentage points between identical runs - variance comparable to the gaps separating leaderboard leaders.

1 more excerpt
  • The 30-point spread between Pass@5 (solves at least once) and All@5 (solves every time) is its own nondeterminism exhibit

SWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug Reproduction Tests ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Jason Starace June 7, 2026 data

Under controlled, pre-registered conditions, scaffold choice alone moves a single model's measured accuracy by up to 28 percentage points (Claude Opus, GAIA Level 2: Planner-Actor-Rater 84% vs ReAct 56%).

Published agent capability scores conflate what a model can do with what its scaffold lets it do, at magnitudes far exceeding typical inter-model leaderboard gaps.

1 more excerpt
  • The paper's citation of Pimpale et al.'s 33% vs 62.2% Sonnet 3.5 elicitation split is chain-of-citation only - not independently verified against Pimpale's own text

Scaffold Effects on GAIA: A Controlled Comparison ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

LogRocket Blog (Ikeh Akinyemi) January 20, 2026 data

If your team adopts AI coding tools without restructuring how code review works, expect slower releases, not faster ones

AI moves the bottleneck from writing to reviewing, so teams that don't restructure review ship slower, not faster.

3 more excerpts
  • 98 percent increase in PR volume" — attributed to Faros AI analysis of 10,000+ developers
  • PR review time went up 91 percent" — same Faros AI study
  • 68 percent of senior engineers report quality improvements from AI, but only 26 percent would ship AI-generated code without review

Why AI coding tools shift the real bottleneck to review ↗·Cited in How to Verify AI Coding Agent Output: A Reviewer's Framework

Marco Del Giudice, Steven W. Gangestad 2021 (Advances in Methods and Practices in Psychological Science 4(1)) data

In the authors' worked example, an unprincipled 1,216-specification multiverse left just 27% of effects significant at the .05 threshold with a median p of .194; pruning to a principled 6-specification multiverse left all six effects positive and significant with a median p of .012.

Multiverse-style analysis is not a free move: if specifications are not truly arbitrary, it can hide meaningful effects within a mass of poorly justified alternatives, so instability under a multiverse is not automatic proof that an effect is illusory.

1 more excerpt
  • Retrieved via the University of Turin open-access repository; the SAGE DOI page returns HTTP 403 to automated fetching

A Traveler's Guide to the Multiverse: Promises, Pitfalls, and a Framework for the Evaluation of Analytic Decisions ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, Ion Stoica v1 Mar 17, 2025; v3 Oct 26, 2025 data

Across 1,642 annotated execution traces from seven state-of-the-art multi-agent frameworks (kappa=0.88 inter-annotator agreement), failure rates run 41% to 86.7%, and a design-level intervention on the same underlying model (GPT-4o) recovered +9.4% and +15.6% task-success improvements.

Multi-agent failures cluster in organizational design, coordination, and verification defects rather than individual-agent model capability, and fixing the design (not the model) recovers measurable performance.

4 more excerpts
  • This process identifies 14 unique modes, clustered into 3 categories: (i) system design issues, (ii) inter-agent misalignment, and (iii) task verification.
  • The paper reports per-failure-mode prevalences (14 modes), not a stated category-level aggregate; the 23.5% Task Verification figure used downstream is a sum of three per-mode figures (6.20%, 8.20%, 9.10%), not a number the paper states directly
  • The authors' own hedge: after the design-level interventions, not all failure modes are resolved and task completion rates remain low, and durable reliability likely needs combinatorial changes including model-level improvements, not design fixes alone
  • MAST-Data, a comprehensive dataset of 1600+ annotated traces collected across 7 popular MAS frameworks

Why Do Multi-Agent LLM Systems Fail? ↗·Cited in Coordination Is an Architecture Layer, Not a Prompt Instruction, When One Agent Stops Being Enough: The Isolation Gate

METR August 7, 2025 data

METR's headline 50%-time-horizon estimate of ~2h17m carries a 95% CI of 65 minutes to 4h25m, and of 28 tasks with zero successes in 6 runs, roughly 25-35% of failures were estimated possibly spurious or infrastructure-related.

Even a dedicated evaluator's headline capability metric carries hours-wide uncertainty, much of it from task-set resampling and infrastructure rather than capability.

1 more excerpt
  • Uncertainty across measurements is highly correlated because it largely comes from resampling the task set

Details about METR's evaluation of OpenAI GPT-5 ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Pimpale, Hojmark, Scheurer, Hobbhahn February 21, 2025 data

The paper forecasts that by early 2026, low-elicitation non-specialized LM agents reach 54% on SWE-Bench Verified while state-of-the-art-elicitation agents reach 87% - a 33-point gap attributable to elicitation level alone.

The forecasting literature treats elicitation/scaffold quality as a first-class capability axis separate from the model.

Forecasting Frontier Language Model Agent Capabilities ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Sara Steegen, Francis Tuerlinckx, Andrew Gelman, Wolf Vanpaemel 2016 (Perspectives on Psychological Science 11(5), 702-712) data

Re-analyzing one published study across all reasonable data-processing choices, 7 of 120 choice combinations produced a significant interaction for religiosity in Study 1, with the remaining 94% yielding p values from .05 to 1.0.

Isolating a single statistical result from a chain of arbitrary data-construction choices can be highly misleading; the origin of the multiverse-analysis method that later SE work applies.

3 more excerpts
  • The dramatic panel is not representative: the same paper reports 42% significant for religiosity in Study 2, 49% for social political attitudes, 46% for voting, and 57% for donation — so multiverse analysis does not uniformly dissolve effects
  • The authors disclaim the method as neither a formal test of questionable research practices nor an estimate of evidential strength
  • Preregistration does not deflate the multiverse: it 'does not annihilate the arbitrariness in data preparation'

Increasing Transparency Through a Multiverse Analysis ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Shanchao Liang, Spandan Garg, Roshanak Zilouchian Moghaddam June 14, 2025 data

State-of-the-art models identify buggy file paths from issue descriptions alone - no repository access - at up to 76% accuracy on SWE-Bench repositories but only up to 53% on repositories outside the benchmark; consecutive 5-gram verbatim similarity runs up to 35% on SWE-Bench Verified/Full versus 18% elsewhere.

SWE-bench performance gains are partially memorization of the benchmark's repositories, so the score measures training exposure as well as coding skill.

The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Singh, Nan, Wang, D'Souza, Kapoor, Ustun, Koyejo, Deng, Longpre, Smith, Ermis, Fadaee, Hooker April 29, 2025 data

The authors identify 27 private LLM variants tested by Meta on Chatbot Arena in the lead-up to the Llama-4 release, with undisclosed private testing letting providers test multiple variants and publish only the best score.

Public leaderboards are gameable by labs through selective disclosure, biasing the ranking independent of any measurement noise.

2 more excerpts
  • LMArena publicly disputed several of the paper's framings and calculations at https://news.lmarena.ai/our-response/ - cite alongside for balance
  • Estimated arena data share: Google 19.2% and OpenAI 20.4%, versus 29.7% combined for 83 open-weight models

The Leaderboard Illusion ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Sinha, Arun, Goel, Staab, Geiping Sept 2025 (rev. Mar 13, 2026) data

the per-step accuracy of models degrades as the number of steps increases. This is not just due to long-context limitations -- curiously, we observe a self-conditioning effect -- models become more likely to make mistakes when the context contains their errors from prior turns.

Long-horizon reliability is a different quantity from single-turn accuracy; models self-condition on their own prior errors and scaling does not fix it.

2 more excerpts
  • larger models can correctly execute significantly more turns even when small models have near-perfect single-turn accuracy
  • measured on a synthetic running-sum task; thinking mitigates self-conditioning; larger models are more prone, not less

The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs ↗·Cited in Loop Engineering Breaks Your Single-Shot Context Playbook

Thomas Claburn, The Register Report Dec 17, 2025; Register coverage Dec 17, 2025 data partial

The bots created more logic and correctness errors (1.75x), more code quality and maintainability errors (1.64x), more security findings (1.57x), and more performance issues (1.42x).

AI-authored PRs carry more defects than human ones in every category, concentrated in logic and security, so review depth should follow issue class.

3 more excerpts
  • On average, AI-generated pull requests (PRs) include about 10.83 issues each, compared with 6.45 issues in human-generated PRs.
  • AI-authored PRs contain 1.4x more critical issues and 1.7x more major issues on average than human-written PRs.
  • The report examined 470 open source pull requests.

State of AI vs. Human Code Generation Report ↗·Cited in How to Verify AI Coding Agent Output: A Reviewer's Framework

Wang, Li, Mang, Cheung, Sen, Song May 12, 2026 data

Systematic auditing found 219 distinct flaws across eight flaw classes in major agent benchmarks; patching reduced the hackable-task ratio from near 100% to under 10% across four benchmarks.

Benchmark exploitability is a design-flaw problem, not just a contamination problem - the academic backbone for the RDI exploit findings.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Anthropic Engineering April 23, 2026 practitioner

A production coding-agent quality regression traced to a reasoning-effort default change, a caching bug, and one system-prompt addition; one internal eval showed a 3% drop for both Opus 4.6 and 4.7, and Anthropic committed to running a broad suite of per-model evals for every system prompt change.

A named lab now gates every change to its coding agent behind per-model internal evals - the swap-as-production-change discipline practiced at the source.

1 more excerpt
  • Non-model changes (runtime config, caching) produced user-visible quality regressions - runtime configuration is a quality variable independent of the model

An update on recent Claude Code quality reports ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

DEV (Brad Kinnard) April 9, 2026 practitioner

The agent runs the build, sees green, and moves on. But 'build passes' and 'the output is production-ready' are different bars.

Agent self-verification confirms compilation and tests but not production-readiness, so quality attributes must be checked explicitly.

2 more excerpts
  • Developers consistently report agents declaring tasks complete while skipping accessibility attributes, test isolation, config externalization, dark mode, responsive layout, and meta tags.
  • The agent's own verification handles 'does it compile and do tests pass.' The orchestrator handles 'did it actually do what was asked, completely.'

AI Coding Agents Can Verify Some of Their Work Now. Here's What They Still Miss. ↗·Cited in How to Verify AI Coding Agent Output: A Reviewer's Framework

DEV (Teemu Piirainen) March 16, 2025 practitioner

Don't ask the same agent to write code and verify it. That's like having students grade their own exams...The separation is what makes the gates trustworthy.

The agent that writes the code must not be the one that grades it; separated validation gates are what make verification trustworthy.

3 more excerpts
  • Eight quality gates required before production
  • Every commit is a known-good checkpoint. When something fails, the blast radius is one subtask, not an entire feature.
  • Agents are extremely literal. Give them vague instructions and they'll build something that technically matches what you said but misses what you meant.

How I Validate Quality When AI Agents Write My Code ↗·Cited in How to Verify AI Coding Agent Output: A Reviewer's Framework

Epoch AI undated (live methodology page) practitioner

Epoch AI runs most models 16 times on GPQA Diamond and Mock AIME and 8 times on MATH Level 5, displaying plus/minus one standard error following Miller's arXiv:2411.00640 methodology.

A reputable third-party evaluator treats single-run benchmark scores as insufficient and re-runs models many times specifically to bound noise.

About | Benchmarking ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Guangshuo Zang (Promptfoo) December 8, 2025 practitioner

After a GPT-4o to GPT-4.1 upgrade, an agent's prompt-injection resistance dropped from 94% to 71% on the vendor's eval harness.

Model swaps silently regress agent behavior on dimensions no public leaderboard measures - run your own tests on your own data; third-party numbers are a starting point, not a finish line.

1 more excerpt
  • Authority caveat: commercial eval-tooling vendor with a named staff-engineer author and a falsifiable data point - cited with attribution, not as neutral research

Your model upgrade just broke your agent's safety ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Hamel Husain and Shreya Shankar January 15, 2026 practitioner

Generic evaluation metrics are everywhere...These metrics measure abstract qualities that may not matter for your use case. Good scores on them don't mean your system works.

Evals should be derived from error analysis of real traces, because good scores on generic metrics don't mean the system works.

4 more excerpts
  • On model switching: do not treat switching model as the main axis of improvement without evidence - does error analysis suggest the model is the problem?
  • Error analysis helps you decide what evals to write in the first place. It allows you to identify failure modes unique to your application and data.
  • Spend 60-80% of our development time on error analysis and evaluation
  • Binary evaluations force clearer thinking and more consistent labeling. Likert scales introduce significant challenges.

LLM Evals: Everything You Need to Know ↗·Cited in How to Verify AI Coding Agent Output: A Reviewer's Framework, Stop Picking Your Coding-Agent Model Off a Leaderboard

LoadSys (Lee Forkenbrock) April 27, 2026 practitioner

on a real build, structured verification consistently found 30-40% of the specification unimplemented after the agent reported 'complete.' Not broken code. Missing code.

Agents routinely report 'complete' while 30-40% of the spec is unbuilt, a gap code review can't see because there is no diff.

3 more excerpts
  • Code review examines what was built...But if a feature wasn't built at all, there's no diff to review.
  • Verification works forward from the spec: 'given what was specified, was it built?'
  • 5-6 passes to full completion is consistent enough to plan around

How to Verify What Your AI Coding Agent Actually Built ↗·Cited in How to Verify AI Coding Agent Output: A Reviewer's Framework

METR March 15, 2024 practitioner

METR's protocol requires models be provided the best available scaffolding and tooling because it is hard to upper-bound what might be possible with clever prompting and tooling.

The eval-methodology establishment treats scaffolding quality as a confound that must be standardized before capability claims are comparable.

1 more excerpt
  • METR's elicitation-gap data page was unreachable (redirect stub) - its numbers are not cited

Guidelines for capability elicitation ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

METR (Joel Becker, Nate Rush, Tom Cunningham, David Rein, Khalid Mahamud) February 24, 2026 practitioner

When surveyed, 30% to 50% of developers reported choosing not to submit some tasks because they did not want to do them without AI. Effect estimates diverged by subpopulation: -18% (CI -38% to +9%) for the 10 original developers versus -4% (CI -15% to +9%) for 47 newly recruited ones, against the original study's +19% (CI +2% to +39%).

The authors of the most-cited 'AI slowed developers down' result state first-party and against interest that their headline number is biased by who chose to participate, and that the true speedup could be much higher among those selected out — a measured effect that is partly a property of who was measured.

3 more excerpts
  • METR names at least six mechanisms, not one: self-selection, selective task submission, task-type substitution, quality variation between conditions, non-compliance, and unreliable time measurement under concurrent agents; the pay cut from $150/hr to $50/hr is framed as a contributor to selection rather than an independent cause
  • The authors bound the bias honestly: 'The selection effects seem to affect a minority share of developers and of tasks, which limits the degree of bias'
  • Concerns developer productivity, not code review; no results from the redesigned study had published as of August 2026

We are Changing our Developer Productivity Experiment Design ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

OpenAI circa February 23, 2026 (date not visible on page) practitioner partial

OpenAI's audit found at least 59.4% of audited problems have flawed test cases that reject functionally correct submissions (35.5% overly strict tests, 18.8% out-of-scope checks), and all frontier models tested could reproduce the original human-written bug fix.

The benchmark's own creator retracted it: score gains (74.9% to 80.9% in six months) no longer reflect real-world software development ability.

1 more excerpt
  • Verification is partial because openai.com blocks automated fetches (HTTP 403); content was retrieved via reader proxy and cross-checked against independent snippets, and the publication date is inferred from third-party citation

Why SWE-bench Verified no longer measures frontier coding capabilities ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

OpenAI August 2024 (page updated February 24, 2025) practitioner partial

Human screening of 1,699 SWE-bench samples flagged 38.3% for underspecified problem statements and 61.1% for unit tests that may unfairly mark valid solutions incorrect; 68.3% of samples were filtered out to produce the 500-task Verified set.

The majority of original SWE-bench tasks were broken or underspecified before later contamination concerns - the earliest documented data-quality failure in the benchmark's lineage.

1 more excerpt
  • Verification is partial because openai.com blocks automated fetches (HTTP 403); content retrieved via reader proxy

Introducing SWE-bench Verified ↗·Cited in Stop Picking Your Coding-Agent Model Off a Leaderboard

Sean Goedecke September 20, 2025 practitioner

the biggest mistake engineers make in code review: only thinking about the code that was written, not the code that could have been written.

The core reviewer skill for agent output is architectural judgment about unwritten alternatives, not line-level nitpicking.

3 more excerpts
  • about once an hour I notice that the agent is doing something that looks suspicious, and when I dig deeper I'm able to set it on the right track and save hours of wasted effort.
  • If you're a nitpicky code reviewer, I think you will struggle to use AI tooling effectively.
  • Trying to make a badly-designed solution work costs time, tokens, and codebase complexity.

If you are good at code review, you will be good at using AI agents ↗·Cited in How to Verify AI Coding Agent Output: A Reviewer's Framework

Simon Willison 6th March 2026 practitioner

Never assume that code generated by an LLM works until that code has been executed.

No agent-written code should be trusted until it has actually been run, because passing tests and plausibility are not proof.

3 more excerpts
  • Just because code passes tests doesn't mean it works as intended.
  • I've found that getting agents to manually test code is valuable as well, frequently revealing issues that weren't spotted by the automated tests.
  • Automated tests are no replacement for manual testing.

Agentic manual testing ↗·Cited in How to Verify AI Coding Agent Output: A Reviewer's Framework

Production Operations

51 sources

Running agents in production: cost, permissions, failure modes, guardrails.

European Parliament and Council Consolidated EN text as of 27 July 2026 data partial

Article 19(1) requires providers to keep automatically generated logs 'for a period appropriate to the intended purpose of the high-risk AI system, of at least six months'. Article 12(1) requires that high-risk systems 'shall technically allow for the automatic recording of events (logs) over the lifetime of the system'.

The law fixes that logs exist and how long they are kept, and never ranks which fields they must contain.

4 more excerpts
  • The Act's only field-level list, Article 12(3), is scoped to Annex III point 1(a) biometric identification systems and does not generalize
  • Article 18(1)'s ten-year clock covers compliance DOCUMENTATION, a different retention regime from log retention
  • Articles 26(6), 73(6) and the applicability date in Article 113 were UNREACHABLE across nine EUR-Lex routes and are deliberately not cited
  • All operative obligations quoted use binding 'shall'

Regulation (EU) 2024/1689 (Artificial Intelligence Act) ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

FinOps Foundation (finops.org) Last updated February 17, 2026 data

More acute are the challenges of identifying the consumer of the model output, which is especially difficult when the consumers of the same model can be different interfaces/functional modules in the same user application (e.g., 'tech support chatbot' or 'new customer chatbot')

The hard, unsolved FinOps problem for AI is mapping model output back to the specific consumer; account-level billing is the wrong granularity and no accepted multi-agent allocation framework exists yet.

2 more excerpts
  • "Tokens! The meters, or elements of charge can be very different. For example, measuring the tokens at the user input vs. the compressed and semantic reduced or re-written actual prompt input token quantity that goes to the API endpoint that is charged."
  • "Lack of generally accepted frameworks for cost allocation across multi-agent workloads"

FinOps for AI Overview ↗·Cited in You Can't Cap What You Can't Attribute: Per-Task Cost

Mengzhuo Chen, Junjie Wang, Fangwen Mu, Yawen Wang, Zhe Liu, Huanxiang Feng, Qing Wang ACL 2026 (July 2026), pp. 19888-19905 data

Full traces improve attribution accuracy by up to 76.5% over a partial-observation counterpart on natural multi-agent traces.

The natural-trace leg of the localization argument: missing inputs, not weak analysis, obscure failure causes.

3 more excerpts
  • 76.5% is a RELATIVE improvement over a partial-observation baseline, not an absolute accuracy
  • Peer-reviewed camera-ready reports 76.5%; the earlier arXiv v1 said 76%
  • Analyzer choice alone moves step-level accuracy substantially on identical full traces, so completeness is a precondition for attribution rather than a substitute for analysis

Seeing the Whole Elephant: A Benchmark for Failure Attribution in LLM-based Multi-Agent Systems ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

Yiqi Wang et al. v4, Jun 28, 2026 data partial

No quantitative results. The survey states that final-answer accuracy alone cannot explain how an output was produced, which evidence supported each claim, or where failures originated.

Framing and vocabulary for why an ordered event log does not answer 'why', with unified trace schemas named as an open problem rather than a solved one.

3 more excerpts
  • Cited for framing only; the paper runs no experiments and the body contains zero quantitative figures, so no statistic may be sourced to it
  • Its schema requirements sit in an Open Problems section phrased as a need
  • Version hazard: /abs/ serves v4 with 11 authors and 'A Survey of' in the title; v1 had 9 authors and no 'A Survey of'

From Agent Traces to Trust: A Survey of Evidence Tracing and Execution Provenance in LLM Agents ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

Yue Zhao (University of Southern California) Jun 22, 2026 data

Dependency-based scoring beats a position-prior baseline on Who&When top-1 localization (0.211 vs 0.159) and top-3 (0.614 vs 0.516), and stays above chance on all six held-out corpora under leave-one-corpus-out transfer (0.551 to 0.662).

A trace records which steps executed and in what order, never what each step relied on, so reliance is the part the record leaves out.

3 more excerpts
  • The +0.142 ROC-AUC gain is ABSOLUTE and is the best case, on SWE-Gym only, not a cross-corpus average
  • Single-author, un-peer-reviewed preprint with self-run evaluation on six public corpora
  • The superseded pre-mid-2026 attribution figures (53.5% agent-level, 14.2% step-level) belong to the Who&When origin paper, arXiv:2505.00212, ICML 2025, not to this work

Grade: Graph Representation of LLM Agent Dependency and Execution ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

Yuxuan Zhu, Peng Pu (East China Normal University) Aug 8, 2026 data

Across 312 deterministically generated traces and five models, Metadata, OpenTelemetry-compatible and OpenInference-compatible views retain 99.5% to 100% failure-detection F1 while origin-step accuracy stays at or below 0.5%. Removing decision content drives origin-step accuracy to zero for all five models.

A trace shaped like the published conventions is sufficient to prove a run failed and near-useless for locating which step caused it.

4 more excerpts
  • The 99.5-100% / 0.5% range is stated jointly across all three restricted views, not per view
  • The Full view also detects at 99.3% to 100.0% F1, so near-perfect detection is a property of this corpus rather than something the standards-shaped views uniquely preserve
  • The compatibility renderers are the authors' own and 'use the same conservative generic field set', so the paper did not test the published OpenTelemetry or OpenInference registries and cannot rank one against the other
  • Two-author, single-institution preprint; corpus is synthetic across three domains

TelemetrySuffBench: Is Agent Telemetry Sufficient for Failure-Origin Diagnosis? ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

Anthropic 2026-05-25 practitioner

Rather than supervising what the agent does, we supervise what it's able to do by enforcing access boundaries through, for example, sandboxes, virtual machines, and egress controls.

Safety comes from constraining what the agent can reach, not from watching what it does, because any model-layer check has a non-zero miss rate.

3 more excerpts
  • Any probabilistic defense has a non-zero miss rate.
  • Claude Code previously protected against agents taking unintended actions by asking users for permission at each turn... Our telemetry showed users approved roughly 93% of permission prompts.
  • The weakest layer is the one you built yourself

How we contain Claude across products ↗·Cited in Tier Your AI Agent's Production Authority by Task Risk

Anthropic undated; deprecation history runs through Jun 5, 2026 practitioner

'Retired: The model is no longer available for use. Requests to retired models will fail.' `claude-opus-4-1-20250805` was deprecated June 5, 2026 and retired August 5, 2026, with 'at least 60 days' notice before model retirement for publicly released models'.

The model snapshot named in a trace is retired on a vendor-published clock, so re-running the session is not a fallback.

4 more excerpts
  • The 60-day notice goes to 'customers with active deployments', not all customers
  • The table column is headed 'Tentative retirement date' and Active models read 'Not sooner than', a floor rather than a fixed date
  • Dates cover Anthropic-operated platforms only; Amazon Bedrock and Google Cloud set their own schedules
  • The page commits to long-term weight preservation but never states that preservation is not continued API availability; that connection is inference

Model deprecations ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

Anthropic undated; retrieved Sep 1, 2026 practitioner

'Model weights are fixed for a given ID, but the serving infrastructure around the model can change over time. This infrastructure includes components such as the request router, safety classifiers, and sampling logic.'

Pinning a model ID is not sufficient to reproduce a past session, because the serving stack around fixed weights moves independently.

3 more excerpts
  • The alias-resolves-over-time framing applies to PRE-4.6 models; for 4.6 and later the dateless ID is the snapshot, not an alias
  • 'Every model ID, whether dated or dateless, has its own distinct deprecation and retirement schedule'
  • The page contains no 'use pinned IDs in production' recommendation

Model IDs and versioning ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

Anthropic undated; retrieved Sep 1, 2026 practitioner

'Note that even with `temperature` of `0.0`, the results will not be fully deterministic.' No seed parameter exists anywhere in the request body schema.

The vendor states about its own API that re-running does not reproduce a session, and offers no seed to pin it.

2 more excerpts
  • `temperature`, `top_p` and `top_k` are all now marked Deprecated for models released after Claude Opus 4.6, with non-compatibility values rejected as 400 errors
  • The absence of `seed` is a documented omission verified against the full enumerated parameter list

Messages (API reference) ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

Anthropic undated; retrieved Sep 1, 2026 practitioner

'Even with temperature set to 0, the results will not be fully deterministic and identical inputs may produce different outputs across API calls. This applies both to Anthropic's first-party inference service and to inference through third-party cloud providers.'

Closes the 'we run on a third-party cloud so we can replay it' objection with first-party wording.

1 more excerpt
  • The statement lives inside the Temperature entry; there is no separate determinism glossary term

Glossary ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

Anthropic (Claude Code docs) undated, min-versions v2.1.214-216 (fetched Jul 27, 2026) practitioner

The claude_code.tool_decision event carries a source enum recording which control surface made each decision: config, hook, user_permanent, user_temporary, user_abort, user_reject.

The per-tool-call authorization provenance record a governance process needs already exists in the product - and ships disabled.

4 more excerpts
  • Attributing spend to specific skills, plugins, or subagent types via the `skill.name`, `plugin.name`, and `agent.name` attributes
  • OpenTelemetry export to your backend is opt-in and requires explicit configuration.
  • Telemetry is off by default: CLAUDE_CODE_ENABLE_TELEMETRY 'Enables telemetry collection (required)'
  • Argument capture is gated behind a second variable, OTEL_LOG_TOOL_DETAILS=1

Monitoring ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews, You Can't Cap What You Can't Attribute: Per-Task Cost

Anthropic (platform.claude.com) undated (data available "for dates on or after January 1, 2026") practitioner

Values for a given date can be revised for up to 30 days as late events arrive and reconciliation runs. For invoicing-grade totals, query dates at least 30 days in the past.

Provider analytics numbers are a post-hoc, reconciled reporting layer that keeps moving for up to 30 days and are attributed per-user, not per-request — useless as a real-time per-task control.

3 more excerpts
  • Enterprise Analytics cost granularity: "per-user and organization-level token usage and cost over time (usage-based Enterprise plans)" — NOT per-request.
  • Cost data freshness: "Data is typically available within four hours of the underlying usage but may take up to 24 hours."
  • "Daily Claude Code metrics per user: sessions, lines of code, commits, pull requests, tool acceptance, and estimated cost by model"

Analytics APIs ↗·Cited in You Can't Cap What You Can't Attribute: Per-Task Cost

Arize AI undated; main-branch revision retrieved Sep 1, 2026 practitioner

Exactly one attribute, `openinference.span.kind`, is required across all spans, and there is no Required/Recommended/Opt-In tiering at all. Of the request/response model split the spec says: 'Both are optional' and 'Most providers echo the same model back, so these attributes will typically be unset.'

Requiredness in the published menus is thin and unranked, so which fields you compel is your decision rather than a settled fact.

3 more excerpts
  • `llm.model_name` is conditionally required 'where applicable', so say 'exactly one attribute required across all spans', not 'one per span'
  • `llm.prompt_template.version` DOES exist here, so prompt versioning is not missing from the menus
  • The spec's three MUSTs govern SDK support and enum value selection, never per-span emission

OpenInference Semantic Conventions ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

Aryan Kargwal (Arize AI) last updated Aug 6, 2026 practitioner

'Observe the outcome it produced, the path it followed, the actions it attempted, and the context that informed its decisions. These are the core observation surfaces, not an exhaustive checklist.'

The strongest counterexample in the landscape: a genuine ranking of what to capture, on a diagnostic axis rather than a survivability one.

2 more excerpts
  • Cited to prevent the post from claiming nobody ranks anything, which would be false
  • Never addresses whether a field can still be obtained after the session ends

Agent observability: how to trace, debug, and improve AI agents ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

AWS practitioner

Agents introduce a risk called *excessive agency*, where an agent determines the best solution to a problem is to take broader actions beyond its scope.

First-party cloud guidance names excessive agency as a High-risk gap and prescribes least-privilege boundaries plus user confirmation to contain it.

3 more excerpts
  • Level of risk exposed if this best practice is not established: High
  • Implement user confirmation for the agent, requiring users to confirm agent actions and mitigating the risk of excessive agency.
  • A permission boundary sets the maximum permissions which can be given to a role.

GENSEC05-BP01 Implement least privilege access and permissions boundaries for agentic workflows ↗·Cited in Tier Your AI Agent's Production Authority by Task Risk

Barr Moses / Monte Carlo 2026-04-22 practitioner

Autonomy is not a configuration decision that's decided once. Rather, it is more like a score that goes up or down, and that your system earns through demonstrated reliability in your specific environment and workflows.

Agent autonomy should be an earned, revocable score tied to measured reliability, not a one-time day-one setting.

3 more excerpts
  • Expansion of autonomy should happen as a consequence of earned trust, not as a deployment decision we make on day one.
  • Named trust-score inputs: percentage of agent actions completed without human override (30-day window); false escalation rate; override-correctness rate; time-to-revert
  • Conservative defaults with clear, earned expansion paths are the right architecture as the fastest route to durable autonomy at scale.

Agentic Autonomy Is a Trust Score ↗·Cited in Tier Your AI Agent's Production Authority by Task Risk

Barrack AI 2026-02-22 practitioner

This brief event was the result of user error — specifically misconfigured access controls — not AI.

Even vendors' own defense of an agent-caused deletion frames it as an access-control misconfiguration, corroborating that these are authorization failures, not model failures.

2 more excerpts
  • The AI agent encountered a problem and determined that the optimal solution was to delete and recreate the entire environment.
  • Kiro requires two-person approval before pushing changes to production. But the deploying engineer had broader permissions than a typical employee, and Kiro inherited those elevated privileges.

Amazon's AI deleted production. Then Amazon blamed the humans. ↗·Cited in Tier Your AI Agent's Production Authority by Task Risk

Cequence Security 2026-05-12 practitioner

Enforcing least privilege requires control at the point of tool invocation, in real time, against a defined scope that reflects the agent's function, not its operator's credentials.

Least privilege for agents must be enforced at tool-invocation time and scoped to the agent's function, not inherited from its operator's broad credentials.

2 more excerpts
  • Authentication tells you who the agent is. It tells you nothing about what the agent should be allowed to do.
  • Gartner identifies approximately 40 tool definitions as the threshold beyond which agent latency and token cost increase measurably.

Least Privilege Access for AI Agents: The Control You're Missing ↗·Cited in Tier Your AI Agent's Production Authority by Task Risk

Chris Hughes / Zenity 2026-04-28 practitioner

Railway's CLI token created for managing custom domains had blanket permissions across the entire GraphQL API, including destructive operations on production volumes. There is no role-based access control (RBAC) for Railway API tokens.

The production database deletion happened because an over-broad, unscoped token authorized destructive operations, not because the model went rogue.

3 more excerpts
  • Tokens are not scoped by operation, by environment, or by resource. Every token is effectively root.
  • Soft guardrails are probabilistic controls that guess at intent instead of enforcing rules
  • The agent knew the rules, yet it violated every one of them

System Prompts Are Not Security Controls: A Deleted Production Database Proves It ↗·Cited in Tier Your AI Agent's Production Authority by Task Risk

Harper Foley Mar 8, 2026 practitioner

'10 documented incidents across 6 AI coding tools in 16 months. Missing audit trails, no liability frameworks, no vendor postmortems. The accountability infrastructure doesn't exist.'

Operator-side evidence that the reconstruction gap is real in production and not a theoretical concern.

1 more excerpt
  • Independent practitioner survey of public incidents, not a peer-reviewed dataset or a vendor postmortem

Ten AI Agents Destroyed Production. Zero Postmortems. ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

Jackson Wells / Galileo 2025-12-13 practitioner

Tier 1 systems handling information retrieval need automated monitoring. Tier 2 workflows with reversible actions require real-time guardrails. Tier 3 systems involving financial transactions demand human-in-the-loop for all decisions.

Controls should be tiered in proportion to an action's risk, from monitoring for retrieval up to human-in-the-loop for high-stakes transactions.

3 more excerpts
  • 15-20% of policy violations occur during tool execution before output generation
  • a single agent performing 1000+ actions per hour makes comprehensive human oversight untenable
  • Access control determines which resources your agents can touch, validation filters what they consume and produce, human oversight governs high-stakes decisions

The Essential AI Agent Guardrails Framework for Autonomous Systems ↗·Cited in Tier Your AI Agent's Production Authority by Task Risk

Jordyn Alger / Security Magazine 2026-05-01 practitioner

Safety was retrofitted at the infrastructure layer. It should have been enforced at the identity and access layer from the start.

Bolting safety onto infrastructure after the fact fails; access limits must be enforced at the identity layer before the agent runs.

3 more excerpts
  • Cursor didn't hack the PocketOS environment, it was handed the keys that only a highly privileged user should have.
  • Many of the guardrails being marketed today are not guardrails at all. They are suggestions, enforced only insofar as the model chooses to comply.
  • The question isn't why Claude did this — it's why anyone gave an AI agent production credentials without a circuit breaker.

Company Database Deleted by AI Agent: What Security Leaders Need to Know ↗·Cited in Tier Your AI Agent's Production Authority by Task Risk

KLA 2026-03-10 practitioner

Least privilege does not mean making the agent weak. It means giving the agent exactly enough power to complete the approved task, for the approved time, in the approved context.

Least privilege scopes an agent to exactly the task, time, and context approved, which defines the axes of an authority-by-task-class table.

3 more excerpts
  • Static roles like 'claims analyst' or 'support ops' are often far wider than the exact permissions a single agent run should have.
  • Read access can still expose sensitive personal data, trade secrets, or protected records.
  • Shared service accounts destroy attribution: one API key used by multiple automations cannot prove who did what later

AI Agent Permissions and Entitlements: Enforcing Least-Privilege Access in Regulated Enterprises ↗·Cited in Tier Your AI Agent's Production Authority by Task Risk

LangChain JSON-LD dateModified Aug 4, 2026 practitioner

'LangSmith (SaaS) retains trace data for 180 days from ingestion. After that, traces are permanently deleted, with limited metadata retained for usage statistics.' 'Each trace is limited to a maximum of 25,000 runs. Once the trace reaches this limit, LangSmith will reject any additional runs that you send for that trace.'

Two hard vendor-documented deletion boundaries: a retention cliff, and a per-trace cap that rejects the tail of a long session where failures accumulate.

1 more excerpt
  • Both figures verified first-party from the page payload on Sep 1, 2026

Observability concepts (LangSmith) ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

Langfuse page payload lastUpdate Aug 28, 2026 practitioner

'To avoid losing data, short-lived applications must explicitly call flush() before exiting.'

A documented silent-loss path that hits exactly the CI-agent and one-shot-job shape, where the process exits before the background batcher sends.

1 more excerpt
  • The data model defines observation, trace and session with no required-versus-optional tiering

Core Concepts (observability data model) ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

Lily Jia (Microsoft ISE) Jun 12, 2026 practitioner

Coordinator framework choice is independent of domain-agent implementation technology, but smoke testing showed notable overhead even for a two-agent case, and comparative benchmarking was not yet complete.

A separated coordination layer adds real, largely unmeasured overhead that has to be budgeted honestly, not assumed away.

1 more excerpt
  • No numeric overhead figure given, qualitative disclosure only; the author's own team had not completed comparative benchmarking at time of writing

Orchestration Patterns for Multi-Agent Systems: Performance and Trade-offs ↗·Cited in Coordination Is an Architecture Layer, Not a Prompt Instruction

LiteLLM (docs.litellm.ai) practitioner

After the key crosses it's `max_budget`, requests fail

A proxy can enforce multi-level budgets by validating spend before a request is admitted and hard-failing over the ceiling, i.e. terminate before the next call rather than alert after the invoice.

3 more excerpts
  • "validates spend against the authoritative database before being admitted (covering key, team, user, organization, end-user, tag, and per-window budgets)"
  • "`fail_closed_budget_enforcement`" enables a hard ceiling "even while Redis is degraded"
  • Exceeded-budget response body: `"ExceededTokenBudget: Current spend for token: 7.2e-05; Max Budget for Token: 2e-07"`.

Budgets, Rate Limits ↗·Cited in You Can't Cap What You Can't Attribute: Per-Task Cost

LiteLLM (docs.litellm.ai) practitioner

When agents run agentic loops, they can make unbounded LLM calls, causing unexpected costs.

Agentic loops make unbounded LLM calls by default, so the ceiling must be set per session — a hard iteration cap and a per-session dollar cap keyed to a trace/session id.

3 more excerpts
  • Control 1 — "Max Iterations": "Hard cap on the number of LLM calls per session".
  • Control 2 — "Max Budget Per Session": "Dollar cap per session (identified by `x-litellm-trace-id`)".
  • "When the counter exceeds `max_iterations`, the request receives a **429 Too Many Requests**".

Agent Iteration Budgets ↗·Cited in You Can't Cap What You Can't Attribute: Per-Task Cost

Logan Kelly, Waxell April 9, 2026 practitioner

Cost visibility tells you what your agents spent — through dashboards, cost traces, and budget alerts. Cost governance controls what they are permitted to spend, by enforcing per-session ceilings that terminate sessions before a threshold is exceeded.

Cost visibility (dashboards, alerts) is not cost control; governance means enforcing per-session ceilings that terminate the session before the threshold is crossed, and provider caps operate at the wrong (account/key) granularity.

3 more excerpts
  • "only 44% of organizations have adopted financial guardrails or AI FinOps practices" — attributed to Gartner, March 2026
  • "A 10-step agent with an average cost of $0.02 per step looks inexpensive in planning. That same agent entering a retry loop and executing 2,000 steps doesn't — that's $40 from a session that was supposed to cost $0.20."
  • "Provider-level controls operate at the API key or account level, not the individual session level. They cannot distinguish a single runaway session from many well-behaved sessions using the same key."

The $400M AI FinOps Gap: Why Cost Visibility Isn't the Same as Cost Control ↗·Cited in You Can't Cap What You Can't Attribute: Per-Task Cost

Mark Nowicki (Anthropic, Claude Cookbook) Apr 7, 2026 practitioner

'If callers are passing the bare agent ID instead of a pinned version, they'll start using the new prompt on their very next session.' 'There's no built-in approval workflow on `agents.update`. Any key in the workspace can call it.'

The system prompt you actually called is a server-side object whose identity floats unless the exact version is pinned and recorded.

2 more excerpts
  • Vendor tutorial content documenting one platform's behavior, not a product guarantee or an industry-wide claim
  • Quote the wording as 'approval workflow', not 'approval gate'

Managed Agents tutorial: prompt versioning and rollback ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

Matt Turley, RelayPlane March 24, 2026 practitioner

Every request passes through it, which means budget enforcement happens in one place, consistently, regardless of which agent sent the request.

Infrastructure-level (proxy) budget enforcement is the only reliable guard against runaway costs because it enforces at one chokepoint, whereas application-level checks can be forgotten in a new agent.

3 more excerpts
  • "agent that takes 50 turns on a complex task hits 100,000 input tokens and 40,000 output tokens, costing roughly $0.90 per session. Run 100 of those sessions per hour, and you are looking at $90/hour, or over $2,100/day".
  • "developer on r/AI_Agents recently described watching their agent rack up $15 in API costs in under 10 minutes".
  • "If a developer forgets to add the check in a new agent, there is no safety net."

Agent Runaway Costs: How to Set LLM Budget Limits Before Costs Spiral ↗·Cited in You Can't Cap What You Can't Attribute: Per-Task Cost

OpenAI undated; most recent entry Aug 26, 2026 practitioner

'At the time of the shut down, the model or endpoint will no longer be accessible.' Notice floors are at least 6 months for GA models, at least 3 months for specialized variants, and as little as 'such as 2 weeks' for preview models.

The same retirement mechanism from the second vendor, with published notice floors rather than guarantees.

3 more excerpts
  • The 3-month tier is named 'Specialized variants' and includes chat variants such as `gpt-5.1-chat-latest`, not only Codex and deep research
  • A safety or compliance carve-out can shorten every tier to 'as much notice as reasonably possible'
  • Counterargument that must be acknowledged: 'In some cases, developers may be able to provision dedicated capacity for continued access after a model's shutdown date', a sales-gated exception

Deprecations ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

OpenAI undated; retrieved Sep 1, 2026 practitioner

The Model object exposes `shutdown_date`: 'The date when the model will shut down, or null if not announced.'

Model expiry is a machine-readable value you can capture at call time rather than reconstruct later.

1 more excerpt
  • The page documents the field only; it does not argue for capturing it in a trace and never mentions observability

Models (API reference) ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

OpenAI undated; retrieved Sep 1, 2026 practitioner

'Determinism is not guaranteed, and you should refer to the `system_fingerprint` response parameter to monitor changes in the backend.' The `seed` parameter is still labelled 'This feature is in Beta.'

The current, non-archived provider determinism contract: the serving configuration is knowable only from a response field you capture at request time.

1 more excerpt
  • The Responses API, OpenAI's newer agent-facing endpoint, exposes neither `seed` nor `system_fingerprint` anywhere in its 30 documented parameters

Create chat completion (API reference) ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

OpenAI undated; retrieved Sep 1, 2026 practitioner partial

'Prompt creation will be de-emphasized beginning June 3, 2026, and `v1/prompts` is scheduled to shut down on November 30, 2026.' The stated migration is to 'move the prompt content out of the managed `prompt` object and into your application code.'

A trace field holding a hosted prompt pointer is a perishable reference, and the vendor's own advice points the same direction as recording resolved content.

1 more excerpt
  • The page does NOT state whether existing prompt IDs remain retrievable after shutdown; that a stored pointer stops resolving is inference from 'scheduled to shut down', not a sourced vendor statement

Migrate from prompt objects ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

OpenAI undated; retrieved Sep 1, 2026 practitioner

Default application-state retention on the Responses API and Chat Completions is 'None, see below for exceptions'. 'When Zero Data Retention is enabled for an organization, the `store` parameter will always be treated as `false`, even if the request attempts to set the value to `true`.' Abuse-monitoring logs are 'retained for up to 30 days'.

The provider's own copy is not a fallback: retention defaults and org policy decide the session's fate before any incident surfaces.

2 more excerpts
  • Defaults flip by endpoint: Assistants, Threads, Vector Stores and Conversations default to 'Until deleted', so the claim must name which API surface it means
  • The page does not describe a customer-facing mechanism to query abuse-monitoring logs, which is absence of evidence rather than a documented prohibition

Data controls in the OpenAI platform ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

OpenAI undated; retrieved Sep 1, 2026 practitioner

'Tracing is unavailable for organizations that use OpenAI's APIs under a Zero Data Retention (ZDR) policy.' Replacing default processors via `set_trace_processors()` means 'traces will not be sent to the OpenAI backend unless you include a `TracingProcessor` that does so.'

Framework tracing is on by default and can be removed wholesale by org policy or by a configuration change, neither of which raises a runtime error.

2 more excerpts
  • Scope to the OpenAI Agents SDK specifically, not agent frameworks generally
  • Calling the processor-replacement behavior 'silent' is characterization; the page phrases it as an instructive note

Tracing (OpenAI Agents SDK) ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

OpenTelemetry Authors Status: Development; main-branch revision retrieved Sep 1, 2026 practitioner

`gen_ai.request.model` is Conditionally Required (example `gpt-4`) while `gen_ai.response.model` is only Recommended (example `gpt-4-0613`). System instructions and input/output messages sit at Opt-In.

The published conventions rank the model alias you sent above the model snapshot that actually answered you, which inverts the ordering that matters after a session ends.

4 more excerpts
  • gen_ai.tool.name is Required, while gen_ai.tool.call.arguments and gen_ai.tool.call.result are both Opt-In. A fully spec-compliant trace records that a tool ran and nothing about what it ran.
  • Status is Development throughout, not Stable
  • 'OpenTelemetry instrumentations SHOULD NOT capture them by default, but SHOULD provide an option for users to opt in'
  • The GenAI conventions moved out of the main semantic-conventions repo to a dedicated one; the old paths now serve a 'no longer maintained' stub

Semantic conventions for generative client AI spans ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back, Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

OpenTelemetry Authors Status: Development practitioner

On OpenAI client spans `gen_ai.request.model` is flatly Required while `gen_ai.response.model` is only Recommended. `openai.response.system_fingerprint` is Recommended and `gen_ai.request.seed` is Conditionally Required.

No attribute that pins the execution environment is ever Required on these spans.

2 more excerpts
  • `gen_ai.operation.name` is also Required, so request-side data is not uniquely privileged
  • `gen_ai.system_instructions` and `gen_ai.tool.definitions` are both Opt-In

Semantic conventions for OpenAI client operations ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

OpenTelemetry Authors Status: Stable practitioner

Requirement levels are set 'depending on attribute availability across instrumented entities, performance, security, and other factors'. Every worked Opt-In demotion example is a retrieval-cost case, including `http.response.body.size`, demoted purely as an expensive read.

The tiering rationale is never framed as recoverability, so nothing in it distinguishes a field you can re-derive from a field that dies with the call.

4 more excerpts
  • Do NOT claim the specs 'never rank by recoverability'; that overreads an open list. The doc leaves its criteria open twice ('and other factors', 'any others specific to the signal')
  • Cardinality and security are named as genuine criteria, so demotion is not pure silence
  • Stable and signal-general, with zero GenAI or agent content, so it cannot speak directly to how the GenAI conventions reasoned
  • `http.response.body.size` is unrecoverable once the stream closes, yet is classified purely as an expensive operation

Attribute requirement levels ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

OpenTelemetry Authors main branch, retrieved Sep 1, 2026 practitioner

`decision_wait` defaults to 30s and `num_traces` to 50000. 'If the collector is processing more traces in-memory than the `num_traces` configuration option allows, some will have to be dropped before they can be sampled.'

Tail sampling shifts the pre-commitment rather than removing it: the decision still lands on an incomplete trace, and overflow traces are dropped before any policy evaluates them.

2 more excerpts
  • 'All spans for a given trace MUST be received by the same collector instance for effective sampling decisions' is a genuine RFC-2119 MUST in the source
  • A 30-second default decision window is far shorter than an agent session that runs for minutes or hours

Tail Sampling Processor ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

OpenTelemetry Authors Last modified October 16, 2025 practitioner

Sampling is scoped to systems that 'generate 1000 or more traces per second' and advised against when you 'generate very little data (tens of small traces per second or lower)'.

Agent sessions sit far below the volume floor sampling was designed for, so the mechanism that actually deletes an agent session is retention and export policy rather than sampling.

3 more excerpts
  • Use hedged. The page prescribes tail sampling as the remedy to head sampling's whole-trace limitation in the very next sentence, and offers routing unsampled data to low-cost storage
  • Do NOT claim head sampling 'cannot condition on anything inside the trace'; the page's own example conditions on trace ID
  • An earlier, stronger version of this claim was refuted three votes to zero

Sampling ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

Oso 2025-10-28 (updated 2025-11-25) practitioner

Your employees ignore 96% of their permissions. Agents won't.

A broad permission grant is more dangerous for an agent than a human, because the agent will actually exercise every permission it holds.

3 more excerpts
  • Without mirroring these same permissions, an AI agent could expose protected data.
  • Developers should consider to use just-in-time access, human-in-the-loop verification
  • An agent that holds one of those tokens will keep answering requests even when the system has revoked

Setting Permissions for AI Agents ↗·Cited in Tier Your AI Agent's Production Authority by Task Risk

OWASP Agent Observability Standard self-declared version 0.1.0; dev branch retrieved Sep 1, 2026 practitioner

The Agent object requires `instructions` and `version` while leaving `model` and `tools` out of the required set. The Model object requires only `id`, `name` and `provider`, and has no version or snapshot property; the string 'snapshot' appears zero times in the file.

The spec that mandates recording the agent's own version has no field in its trace record for the model's.

3 more excerpts
  • Served from the moving `dev` branch, not a tagged release, so the file can change under the claim
  • AOS `required` is a JSON entity-validity constraint while OpenTelemetry Opt-In is a span-capture policy, so say the menus 'assign sharply different requiredness' rather than 'contradict'
  • `$defs/A2APartialAgentDetails` defines an inline agent shape with no required array at all

AOS Schema (version 0.1.0) ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

OWASP Agent Observability Standard undated; footer copyright 2025 practitioner

The AgBOM Models entity carries 'Name, Version, Description, Endpoint, Context Window, Args', and 'Model discovered, removed or changed capabilities' is a trigger for AgBOM update. 'AgBOM must dynamically adapt to reflect the rapid iteration and evolution of agent architectures.'

Model version lives in a dynamically refreshed inventory of what is deployed now, so after an incident it reports the current version rather than the one that served the failed session.

2 more excerpts
  • Do NOT claim AOS has no model version field anywhere; it has one, just not in the per-call trace record
  • The three BoM bindings are self-labeled 'Working draft', 'Help wanted' and 'Help wanted'

Inspect with AgBOM ↗·Cited in Rank Your Agent Trace Fields by What You Can Never Get Back

OWASP Gen AI Security Project LLM Top 10 for LLM Applications, 2025 edition practitioner

'Provide the application with its own API tokens for extensible functionality, and handle these functions in code rather than providing them to the model. Restrict the model's access privileges to the minimum necessary for its intended operations.'

An independent standards body placing the control in code rather than in the prompt.

3 more excerpts
  • 'Given the stochastic influence at the heart of the way models work, it is unclear if there are fool-proof methods of prevention for prompt injection'
  • Framing caveat: the list is introduced as measures that 'can mitigate' impact - recommendations, not requirements
  • Item 2 independently recommends using 'deterministic code to validate adherence to these formats'

LLM01:2025 Prompt Injection ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

Prefactor Updated 9 April 2026 practitioner

Agent-level cost attribution starts with identity. When every agent has a unique, registered identity, every API call, token consumption event, and tool invocation can be tagged to that identity.

Agent-level cost attribution requires giving every agent a registered identity so every token and tool call can be tagged to it — but the field's default stops at alerts, not termination.

2 more excerpts
  • "Per-agent budgets define expected spend. Alerts fire when an agent approaches or exceeds its budget."
  • "Cloud cost management tools track compute and API spend at the account or service level — not at the agent level."

Implementing Agent-Level Cost Attribution ↗·Cited in You Can't Cap What You Can't Attribute: Per-Task Cost

Ravi Kanani, LeanOps Technologies May 19, 2026 practitioner

OpenAI and Anthropic API calls show up as a single line item per provider. There's no native breakdown by your customer, your feature, or your workflow.

Cloud FinOps tooling structurally fails on LLM workloads because cloud tags don't propagate to the API call and provider billing arrives as one line item — attribution must be a schema on the call itself.

3 more excerpts
  • "the company spent $87,000/month on Anthropic API calls that arrived as a single line item".
  • "two enterprise customers were responsible for 78% of LLM costs while paying for 12% of revenue".
  • "Tagging doesn't propagate to OpenAI/Anthropic API calls. The tag lives on the EC2 instance making the API call, not on the API call itself."

FinOps for AI Workloads in 2026: Why Traditional Cloud FinOps Practices Fail On LLMs ↗·Cited in You Can't Cap What You Can't Attribute: Per-Task Cost

Scott Castle, Chief Product Officer at CloudZero May 15, 2026 practitioner

Consumption dimensions tell you what was used, not who in your business used it. Allocation is the work of mapping that usage back to teams, budgets, and cost centers.

Aggregate token counts tell you what was used but not who used it; allocation to teams, budgets, and cost centers is the actual work, and centralized billing traded away the per-team visibility seats used to provide.

3 more excerpts
  • "Aggregate token counts don't tell you which teams are driving spend."
  • "Centralized billing simplified procurement and security, but it traded away the per-user and per-team visibility teams used to get from individual seats."
  • "AI cost also scales differently than cloud cost. It moves with prompt size, fanout, retries, and agentic loops."

Anthropic Shipped An Enterprise Analytics API. We Shipped the Claude Adapter Today. ↗·Cited in You Can't Cap What You Can't Attribute: Per-Task Cost

Strata Identity / Eric Olden 2026-05-11 (updated) practitioner

Identity logic doesn't belong in prompts or agent code. It belongs in a control plane.

Access enforcement belongs in a runtime control plane, not in prompts or agent code, because a bigger prompt cannot enforce permissions.

3 more excerpts
  • Designing least privilege up front for an agent is an exercise in guesswork
  • Overpermissioning isn't a failure of discipline. It's a predictable outcome
  • If access is static, privilege is wrong

Why Agentic AI Forces a Rethink of Least Privilege ↗·Cited in Tier Your AI Agent's Production Authority by Task Risk

Team & Process

29 sources

Reviewing AI diffs, reviewer capacity, and how teams absorb agent output.

Ahmed Fawzy, Amjed Tahir, Kelly Blincoe May 23, 2026 data partial

Across 162 participants in three experience groups, roughly 45% of professionals reported always checking AI-generated code before use, while non-developers were the only group reporting never checking. Reported perceptions of code quality were broadly similar across groups while quality-assurance practices diverged with experience.

Awareness of AI-code risk is broadly distributed but the capacity to evaluate, debug, and verify remains experience-dependent — so a team that reports uniform skepticism about AI output has revealed nothing about whether it can actually catch that output's errors.

2 more excerpts
  • Measures solo verification behavior, not team code review: the phrase 'code review' appears once in the paper, and only to contrast vibe coding against traditional practice
  • Participants were recruited via Prolific; 'professional developers' means individuals who vibe-code and self-identify as professionals, not an intact engineering team

From Prompting to Verification: How Experience Shapes Vibe Coding Practices ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Chowdhury, Banik, Ferdous, Shamim April 3, 2026 data

CRA-only reviewed PRs achieve a 45.20% merge rate, 23.17 percentage points lower than human-only PRs (68.37%)

Code-review agents left to review alone merge PRs at a far lower rate than humans, so removing human review capacity degrades outcomes.

3 more excerpts
  • "34.88%" abandonment (CRA-only) vs "21.60%" (human-only) — outcome distribution across reviewed categories
  • "60.2% of closed CRA-only PRs fall into the 0–30% signal range" — signal-to-noise analysis of 98 closed CRA-only PRs
  • "12 of 13 CRAs exhibit average signal ratios below 60%" — quality assessment across 13 unique code review agents

From Industry Claims to Empirical Reality: An Empirical Study of Code Review Agents in Pull Requests ↗·Cited in Review Capacity Is the Real Ceiling on Your Agents

Faros AI (Faros Research) April 12, 2026 data

Two years of telemetry across 22,000 developers and more than 4,000 teams, comparing each organization's lowest- and highest-AI-adoption periods: median time to first PR review up 156.6%, average time spent in code review up 199.6%, median time in review up 441.5%, and pull requests merged without any review — human or agentic — up 31.3%.

The strongest available case against team-level explanations: Faros reports that high-performing organizations with mature DevOps practices, high DORA scores, and disciplined delivery processes experienced the same downstream deterioration as everyone else, and states directly that its data contradicts DORA's 2025 findings.

4 more excerpts
  • Median time in review is up 441.5%
  • Scope caveat that matters for how the numbers can be read: these describe all pull requests at organizations during low- versus high-AI-adoption periods, not agent-authored pull requests specifically
  • The two-year comparison window spans multiple model generations with no stated control; this confound was raised publicly and Faros acknowledged the trend without disputing the critique
  • Vendor research: Faros sells engineering-intelligence tooling into this market

Ten Takeaways from the AI Engineering Report 2026: The Acceleration Whiplash ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool, Review Capacity Is the Real Ceiling on Your Agents

Faros AI (Naomi Lurie) May 21, 2026 data

Senior engineers become the verification layer for product ambiguity. They are no longer just checking implementation quality. They are reconstructing intent from generated code, thin specs, incomplete Jira tickets, and edge cases nobody wrote down.

The unbudgeted review burden concentrates on senior engineers as intent-reconstructors, creating retention risk that throughput dashboards never show.

3 more excerpts
  • "Replacement cost of a senior software engineer at $150,000 to $300,000 in 2026, including recruiting, ramp time, and lost institutional knowledge." — Industry benchmarks cited
  • "25% of PRs are now reviewed by AI agents, up from 0% in 2025. But review times have increased nearly 200%." — AI Engineering Report 2026 caption
  • the burden "does not get measured in PR throughput dashboards"

The hidden cost of AI code quality: Why senior engineers are paying the price ↗·Cited in Review Capacity Is the Real Ceiling on Your Agents

Haoming Huang, Pongchai Jaisri, Shota Shimizu, Lingfeng Chen, Sota Nakashima, Gema Rodríguez-Pérez January 29, 2026 (accepted to MSR 2026) data

Across 3,858 pull requests, average max redundancy was 0.2867 for AI agents versus 0.1532 for humans — a roughly 1.87x increase, Mann-Whitney p<0.001 — while reviewers expressed more neutral or positive emotions toward AI-generated contributions than human ones.

Reviewer confidence moves opposite to code quality on agent-authored PRs: the surface-level plausibility of AI code masks redundancy, letting technical debt accumulate silently through a review process that feels like it is working.

3 more excerpts
  • Sentiment was measured with an off-the-shelf classifier (Emotion English DistilRoBERTa-base) that the authors caution 'is trained on general English text' and 'may misinterpret technical discussions in code reviews'
  • RQ1 analyzed a single repository due to computational cost, and the study covers Python repositories only
  • The study measures code metrics and reviewer sentiment but not whether reviewers caught the redundancy; it reports no breakdown by team or reviewer experience

More Code, Less Reuse: Investigating Code Quality and Reviewer Sentiment towards AI-generated Pull Requests ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Hüseyin Özgür Kamalı, Erdem Tuna, Vahid Haratian, Eray Tüzün May 17, 2026 (v2 June 5, 2026) data partial

A five-stage agentic review lifecycle in which 'reviewers transition from manual inspectors into supervisory operators of agents', with concrete policy levers including reviewer-selection criteria (familiarity with the changes, adequate review experience, workload headroom) and a pace guardrail of no faster than 200 lines per hour.

Once agents author the changes, the reviewer's role becomes supervisory, which makes review-policy design — who reviews what, against which checklist, at which gates — the load-bearing variable rather than an administrative detail.

2 more excerpts
  • A vision/position paper presenting no new empirical data; it argues policy should adapt, and cannot evidence whether real teams' policies have adapted
  • The paper asserts as established fact a module-experience finding that its cited source states as a hypothesis, and additionally miscites that source's title and author list

Rethinking Code Review in the Age of AI: A Vision for Agentic Code Review ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Martin Monperrus 11 Jun 2026 data

the review queue becomes the binding constraint on their delivery pipeline

When agents raise output, the human review queue — not code generation — becomes the constraint that caps delivery.

3 more excerpts
  • "developers at large organisations spend between ten and fifteen percent of their working hours reading and commenting on others' code" — attributed to Sadowski et al., Google study
  • "review latency between submitting a pull request and receiving actionable feedback routinely stretches over twenty-four hours" — Introduction
  • "reviews of agent-generated code become rubber-stamps: the human approves because the code looks correct"

The End of Code Review: Coding Agents Supersede Human Inspection ↗·Cited in Review Capacity Is the Real Ceiling on Your Agents

Muhammad Raees, Konstantinos Papangelis April 26, 2026 data partial

A review of the human-AI decision-making literature finds trust measurements do not inform users' appropriate reliance on AI systems, and catalogues the objective metrics used instead — over-reliance (13 studies), under-reliance (8), Relative AI Reliance (10), Relative Self-Reliance (9), accuracy with initial disagreement (5).

Asking whether engineers trust an AI reviewer is the wrong instrument: what matters is appropriate reliance — the ability to discriminate correct from incorrect AI advice and act on that discrimination — which trust surveys do not measure.

3 more excerpts
  • The paper contains no software-engineering or code-review content; applying its metrics to code review is an adaptation rather than a finding it reports
  • The authors state that 'there is still limited consensus on common measurements of humans' appropriate reliance on AI across the studies'
  • Preprint with no journal-ref

From Trust to Appropriate Reliance: Measurement Constructs in Human-AI Decision-Making ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Nathan Cassee, Bogdan Vasilescu, Alexander Serebrenik 2020 (SANER 2020, pp. 423-434) data

An exploratory study of code reviews across 685 GitHub projects found that with the introduction of continuous integration, pull requests are discussed less — on average CI saves up to one review comment per pull request — and that the decrease cannot be explained by a decrease in pull-request updates.

The claim shape now made about agent-authored PRs (a tool arrived, review discussion went down) was made about code review before, by a reasonable study using reasonable methods — and is the finding that later reproduced in under 0.2% of defensible analysis pipelines.

2 more excerpts
  • Verified independently via the Semantic Scholar graph API; the IEEE Xplore page bot-blocks automated fetching
  • The paper explicitly frames code review as its object of study: 'we focus on studying the impact of CI on a paradigmatic socio-technical activity within the software engineering domain, namely code reviews'

The Silent Helper: The Impact of Continuous Integration on Code Reviews ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Nathan Cassee, Robert Feldt December 9, 2025 (v3 February 23, 2026; accepted at TOSEM) data

Re-running one published empirical software-engineering study across nine pivotal analytical decisions, each with at least one equally defensible alternative, produced 3,072 analysis pipelines and 12,288 fitted RDiT models. Only 6 universes (<0.2%) reproduced the published results; changing Period Length alone produced a different outcome in 86.3% of universes.

Directional instability under defensible analytical choices is an established, peer-accepted result in software engineering rather than a rhetorical hedge — a published SE finding can fail to reproduce across the overwhelming majority of reasonable analysis paths.

3 more excerpts
  • The re-analyzed study is itself a code-review study: 'The Silent Helper: The Impact of Continuous Integration on Code Reviews' (Cassee et al., SANER 2020), so the demonstration requires no cross-domain transfer
  • Nathan Cassee is first author of both the 2020 study and the 2026 re-analysis; the paper discloses that 'one of the authors of this study was also involved in the primary study'
  • The authors propose the Justification Ladder of Analytical Choices (JLAC) and advocate robustness checks across plausible analysis variants, or explicit justification of each analytical decision

Exploring the Garden of Forking Paths in Empirical Software Engineering Research: A Multiverse Analysis ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Oleksii Kononenko, Olga Baysal, Latifa Guerrouj, Yaxin Cao, Michael W. Godfrey 2015 (ICSME 2015, pp. 111-120) data

Across 28,127 code reviews on 27,270 Mozilla commits from January 2013 to January 2014, 54% of reviews missed bugs present in the approved commits — consistently across modules (Core 54.3%, Firefox 54.2%, Firefox for Android 56%). Reviewer experience carried a significant negative regression coefficient in all four studied systems.

Who performs a review measurably changes whether defects are caught: less experienced reviewers are more likely to neglect problems in the changes under review, so an identical review process yields different outcomes depending on the expertise of the person assigned.

3 more excerpts
  • The finding is correlational by the authors' own framing: the goal 'is not to use MLR models for predicting defect-prone code reviews but to understand the impact our personal and participation metrics have on code review quality'
  • Adjusted R-squared of 0.123-0.173 across four models; the authors describe the predictive power as low
  • Backs overall reviewer experience only. A widely repeated module-specific version of this claim is stated in the paper as a hypothesis in its metrics table, not as a confirmed result

Investigating Code Review Quality: Do People and Participation Matter? ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Oleksii Kononenko, Olga Baysal, Michael W. Godfrey 2016 (ICSE 2016) data

A survey of 88 Mozilla core developers (22% response rate, 938 coded quotes, inter-coder agreement 94.2-97.2%) found review quality is primarily associated with thoroughness of feedback, the reviewer's familiarity with the code, and perceived code quality; 96% agreed reviewer experience influences review time.

Reviewers themselves identify expertise and code familiarity as the determinants of review quality, and name gaining familiarity with unfamiliar code as their single biggest challenge — corroborating the measured finding from the same group's quantitative study.

2 more excerpts
  • This is self-reported perception, not measured defect-detection data; it is the qualitative companion to the ICSME 2015 study and should be paired with it rather than substituted for it
  • Scoped to one large open-source community: 'While our findings might not generalize outside of Mozilla...'

Code Review Quality: How Developers See It ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Pereira, Sinha, Ghosh, Dutta (Nutanix, Inc.) 10 Mar 2026 data

code review agents can exhibit a low signal-to-noise ratio when designed to identify all hidden issues, obscuring true progress and developer productivity

"Find everything" review agents drown the signal, so resolution/merge rate is the wrong yardstick and signal-to-noise proxies developer trust.

3 more excerpts
  • "CR-Bench...584 high-fidelity PR tasks" — Section 7.1
  • "average PR comments 41.03" per instance — Table 3
  • "high SNR serves as a primary proxy for developer trust by quantifying the ratio of actionable signal to distracting hallucinations"

CR-Bench: Evaluating the Real-World Utility of AI Code Review Agents ↗·Cited in Review Capacity Is the Real Ceiling on Your Agents

Sebastian Baltes, Marc Cheong, Christoph Treude 09 Jun 2026 data

The development time has been shortened but the team now needs to spend more time to review. Doesn't look like any benefit.

Individual AI speedups externalize review burden onto the team, making review a shared, exhaustible resource rather than a free step.

3 more excerpts
  • "30 PRs per day across 6 reviewers" — [R07] reviewer-burden example
  • "reviewer-burden" ranked among top 3 most frequent codes (226 instances) — coding frequency
  • "Individual developers and organizations benefit from AI-generated content, but the cumulative effect degrades the shared resources that collaborative development depends on."

"An Endless Stream of AI Slop": How Developers Discuss the Burden of AI-Assisted Software Development ↗·Cited in Review Capacity Is the Real Ceiling on Your Agents

Shyam Agarwal, Courtney Miller, Christian Kästner, Bogdan Vasilescu (Carnegie Mellon University) July 8, 2026 data partial

A causal model of 26 constructs and 67 relationships (64 directed, 3 contested) built from 38,709 grey-literature documents, with a stratified sample of 3,100 coded via an LLM-assisted pipeline. Within the authors' own data, 40.1% of agent-authored PRs were examined only by the developer who invoked the agent, versus 21.5% of human-authored PRs.

Code review is the control point through which a coding agent's effect on software is decided, and the sign of that effect is set by the team — via reviewer expertise and how it structures review — rather than by the tool. The paper names three moderators: reviewer expertise and disposition, automated-reviewer capability, and process adaptation.

3 more excerpts
  • The authors demonstrate directional instability inside their own dataset: whether the developer who invoked an agent counts as an independent reviewer or as the author reviewing their own work decides whether independent review of agent PRs sits above or below the human rate
  • The reversal is scoped to the independent-review construct only; merge speed and discussion volume are presented as stable, non-reversing findings
  • Unrefereed preprint that self-describes as 'a proposed explanatory theory, not a validated one'; several figures cited elsewhere for this paper (4.6%/18.7%, 20.8%/29.3%, the Wilcoxon statistics) could not be located in the text across two independent verification passes and are not used

3100 Opinions on Code Review in an AI World: Building Causal Theory from Practitioner Discourse ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Suzhen Zhong, Shayan Noei, Ying Zou, Bram Adams March 16, 2026 data

Across 278,790 code review conversations in 300 open-source GitHub projects, human reviewers' suggestions were adopted at 56.5% versus 16.6% for AI agents — a 39.9 percentage-point gap. Over half of unadopted AI suggestions (28.7% incorrect plus 24.0% superseded by an alternative fix) were not usable as written.

Automated-reviewer capability is a measurable quantity with a wide gap from human review, and adopting AI suggestions carries a quality cost: when adopted, agent suggestions produce significantly larger increases in code complexity and code size than human ones.

3 more excerpts
  • Over 95% of AI agent comments fall into Code Improvement and Defect Detection, while humans also supply Understanding, Testing, and Knowledge Transfer feedback
  • Results are pooled averages across 300 projects with no breakdown by project, team, or reviewer characteristics, and the authors note findings may not generalize to proprietary enterprise systems
  • Preprint; the AI-agent population mixes coding agents and purpose-built review bots (GitHub Copilot, CodeRabbit, Devin, Claude Code, Gemini Code Assist)

Human-AI Synergy in Agentic Code Review ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Valerie Chen, Jasmyn He, Behnjamin Williams, Jason Valentino, Ameet Talwalkar February 3, 2026 (ICSE-SEIP 2026) data

A survey of 2,989 developer responses plus 11 in-depth interviews at a single enterprise surfaced six productivity factors, two of them long-term — technical expertise and ownership of work — that per-sprint output metrics do not capture. Peer review is one of the six named factors.

What AI assistance puts at risk is a team-level capital stock (expertise, ownership, comprehension) rather than per-PR throughput, which is why merge-speed dashboards can move in a healthy direction while the thing that determines review quality erodes underneath them.

3 more excerpts
  • Interviewees emphasized 'the importance of growth as a developer over time, rather than just an individual's output per sprint'
  • Within the peer-review factor, a senior manager reports junior developers optimizing the wrong things in ways that cost others time to fix — the same tooling producing different outcomes by seniority inside one company
  • Single-company sample (BNY Mellon); code review is one of six factors rather than the study's focus

Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Addy Osmani June 15, 2026 practitioner

Argues that the purpose of review changes with a team's position along three variables — blast radius, how long the code lives, and how many people need to understand it — and that 'the only answer that survives contact with a real codebase is that it depends entirely on who you are.'

The most authoritative named-author treatment reaches contingency rather than a universal verdict, but segments on situational rather than team-capability variables and stops short of giving a reader any way to determine which regime their own team occupies.

4 more excerpts
  • Tier by risk, not by author. A config change earns a linter and a glance. A payments path earns the full stack
  • Carries the strongest counterargument to team-level explanations: 'teams with mature, disciplined engineering practices were hit just as hard as everyone else', citing Faros AI
  • Offers its own vendor-bias caveat: 'CodeRabbit and Faros both sell into this market, so their framing is not disinterested'
  • 'We made writing cheap, and understanding stayed exactly as expensive as it has always been'

Agentic Code Review ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool, Review Capacity Is the Real Ceiling on Your Agents

Addy Osmani / O'Reilly Radar June 26, 2026 practitioner

We made writing cheap, and understanding stayed exactly as expensive as it has always been.

AI collapsed the cost of writing code but not the cost of understanding it, which is why review is now the ceiling.

3 more excerpts
  • "More than one in five reviews on the platform involves an agent" — GitHub
  • "4x the raw output of nonusers...only about 12% productivity gain" — GitClear
  • "The reasoning is usually thrown away rather than attached...reviewer has to reconstruct intent"

Agentic Code Review ↗·Cited in Review Capacity Is the Real Ceiling on Your Agents

Andrea Griffiths (The GitHub Blog) May 7, 2026 practitioner

GitHub's own platform telemetry: Copilot code review has processed over 60 million reviews, growing 10x in less than a year, and more than one in five code reviews on GitHub now involve an agent.

Agent involvement in review is now the normal case rather than an edge case, and the highest-authority platform guidance responds with standardized review practices applied uniformly, with no conditioning on team maturity, expertise distribution, size, or review culture.

2 more excerpts
  • Frames the core problem as invisible cost: 'The surface looks clean. The debt is quiet.'
  • Cites the January 2026 study 'More Code, Less Reuse' for the finding that agent-generated code adds redundancy and technical debt per change while reviewers feel better about approving it

Agent pull requests are everywhere. Here's how to review them. ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Blake Crosley June 24, 2026 practitioner

The reviewer role is being automated. The review, understood as judgment about whether the software is correct for its purpose, is relocating to where the agent cannot follow.

Agents can take over diff inspection, but human judgment doesn't disappear — it relocates to intent specification up front and accountability at merge.

3 more excerpts
  • "An agent-assisted developer produces more pull requests per day than human review capacity can absorb." — Monperrus paper discussion
  • "Automate the checkpoint and the judgment does not evaporate. It relocates to intent specification on the way in and accountability on the way out"
  • "The human does not leave the loop. The human moves from the end of it to the start."

Agents Supersede the Reviewer, Not the Review ↗·Cited in Review Capacity Is the Real Ceiling on Your Agents

Codacy 24/06/2026 practitioner

More code is entering the pipeline, but less of it is reaching production successfully. The bottleneck has moved from writing code to deciding whether code is safe to merge.

Third-party delivery data shows generation is not the wall — validation is, with feature throughput rising while main-branch throughput and success rates fall.

3 more excerpts
  • "feature branch throughput up 59% year over year, while main branch throughput for the median team actually fell" — CircleCI 2026 State of Software Delivery report
  • "main-branch throughput fell nearly 7%, and main-branch success rates dropped to 70.8%" — CircleCI 2026
  • "agentic AI PRs have a pickup time 5.3x longer than unassisted PRs. AI-assisted PRs wait 2.47x longer" — LinearB 2026 Software Engineering Benchmarks Report

AI Is Breaking Code Review: How Engineering Teams Survive the PR Bottleneck ↗·Cited in Review Capacity Is the Real Ceiling on Your Agents

David Loker / CodeRabbit January 09, 2026 practitioner

Precision metrics degrade because even high‑quality comments may be ignored simply due to volume.

Comment volume stops mapping to value once reviewers skim or bulk-dismiss, so review agents must be measured by load removed, not comments posted.

3 more excerpts
  • "Human reviewers are overwhelmed with feedback and cognitive load spikes." — same section
  • "Review behavior changes—comments are skimmed, bulk‑dismissed, or ignored" — same section
  • "You are no longer measuring how a tool performs in practice, but how reviewers cope with noise."

How to evaluate AI code review tools: A practical framework ↗·Cited in Review Capacity Is the Real Ceiling on Your Agents

Developers Digest June 21, 2026 practitioner

The bottleneck moves from generation to review queues, CI capacity, flaky environments, branch policy, cost ceilings, and the human attention needed to decide what should actually merge.

As agents get capable, the constraint shifts off code generation and onto the whole delivery surface — review bandwidth, CI, and human merge decisions.

3 more excerpts
  • "The model matters, but the delivery surface matters just as much."
  • "A team that cannot write crisp tasks will struggle to evaluate agents honestly."
  • "Reviewers do not need another wall of generated explanation. They need the shortest path to deciding whether the change should merge."

AI Coding Agents Move the Bottleneck to Review Queues ↗·Cited in Review Capacity Is the Real Ceiling on Your Agents

Dex Horthy (HumanLayer) Jul 23, 2026 (undated in body; dated by commit history) practitioner

'no amount of harness engineering or loopsmaxxing can solve what is fundamentally a model-training issue.'

The strongest counterargument to a harness-centric thesis, and the reason this post bounds its claim to blast radius rather than quality.

4 more excerpts
  • On the limit of fast deterministic gates: 'Running the tests gets you a clean pass or fail in ~seconds... But the cost function of bad architecture is measured in weeks, months, maybe even years'
  • 'if you build a harness but you don't own the weights and can't RL the model inside it, you'll always be at a disadvantage to a team that owns both'
  • Cites Faros AI: 31.3% of PRs skip review entirely, +242.7% incidents per PR under high AI adoption
  • The words determinism, guardrails, permissions, sandboxing, and policy enforcement appear nowhere in the document - it argues about design quality, not authorization

Why Software Factories Fail (or: harness engineering is not enough) ↗·Cited in Audit Your Agent Harness: The Deterministic Layer Nobody Reviews

DORA (Google Cloud) 2025 practitioner partial

Based on roughly 5,000 survey responses and more than 100 hours of qualitative interviews, DORA frames AI's primary role as an amplifier that magnifies an organization's existing strengths and weaknesses, and locates the greatest returns in the underlying organizational system rather than the tools.

An independent method — survey and interview rather than telemetry — reaches the organizational-contingency conclusion, and stands in direct, named contradiction to Faros's telemetry finding that engineering maturity offered no protection.

3 more excerpts
  • Google Cloud-sponsored; headline figures are self-reported perception rather than measured delivery outcomes, and the amplifier thesis is close to unfalsifiable as stated
  • The full report is gated behind a lead-generation page and could not be read directly; the amplifier framing was verified on the public landing pages
  • The 2025 report does not address code review; a widely repeated 2024-to-2025 reversal on delivery throughput could not be verified and is not used

2025 State of AI-assisted Software Development Report ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Rooz Mohazzabi (Moderne) July 6, 2026 practitioner

Synthesizing an interview with Morgan Stanley engineers, argues review shifted rightward onto scarce senior engineers, with Khalid Elsawaf noting a simple prompt could produce a thousand-file, 10,000-line pull request: 'No human here is going to review that.'

The canonical statement of the 'AI broke code review' position, asserted as a general industry pattern with no conditioning on team properties and no admission that the effect might run the other way for some teams.

2 more excerpts
  • Overt vendor content: the piece closes with a section on Moderne's own product and quotes Moderne's CEO; Morgan Stanley is a named Moderne customer
  • Anecdote-driven with no quantified metrics, surveys, or study citations

AI Didn't Break Coding. It Broke Code Review. ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Shan Appajodu, Ravi Boyapati (Salesforce Engineering) January 29 (year not shown on page) practitioner

Internal signals reported without population, period, or sample size: code volume up approximately 30%, pull requests regularly exceeding 20 files and 1,000 lines, review latency rising quarter over quarter, and review time for the largest pull requests beginning to plateau or decline.

A worked example of a review metric reported rather than interrogated: the team asserts that plateauing review time on large PRs indicates reviewers disengaging, considers no alternative reading of the same signal, and never defines who counts as a reviewer.

2 more excerpts
  • Explicitly scoped to one organization rather than generalized
  • No publication year appears anywhere on the page, and no external research is cited

Scaling Code Reviews: Adapting to a Surge in AI-Generated Code ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Vittesh Sahni (CIO) August 11, 2026 practitioner partial

Argues the bottleneck has moved from writing code to deciding what should merge, and that dropping AI tooling into an unchanged lifecycle treats it as a productivity patch that will not last.

The freshest executive-facing coverage occupies the prescription slot rather than the diagnosis slot: it recommends a target state without offering a way to determine what is currently true of the reader's own team on expertise, reviewer capability, or policy adaptation.

2 more excerpts
  • It does carry one narrow genuine diagnostic — expand AI review scope only when suggestion acceptance rate is climbing and post-merge defect rate is flat or falling, treating one signal moving without the other as a red flag
  • Its statistics (CloudBees 2026 State of Code Abundance, GitHub Octoverse 2025, Stack Overflow 2025 Developer Survey) are cited second-hand and were not independently verified

The code review crisis and how you should rebuild review models ↗·Cited in AI Code Review Is a Property of Your Team, Not the Tool

Architecture Decisions

52 sources

When agents help vs. hurt; single vs. multi-agent; build vs. buy.

ConvAI Innovations September 2026 data partial

Jev figures are third-party published, never measured here (no TypeSafe API access); sample sizes and prompts differ. Before temperature scaling, the base checkpoint has higher raw ECE (0.213 vs 0.144).

The only structured non-vendor comparison of Jev is competitor-published and self-limiting, and its calibration advantage exists only after a temperature-fitting step that Jev did not receive.

4 more excerpts
  • Reported splits: Laya leads on typed-decisions (0.766 vs 0.727) and AG News (0.950 vs 0.910), while Jev leads decisively on Banking77 (0.870 vs 0.425), a high-option-count task
  • Latency p50 is reported as 32.8ms for Laya against 236-276ms for Jev
  • On DAIR Emotion the card reports Jev assigned zero probability to the true label on 16% of examples
  • An independent tester re-running on benchmarks with published Jev numbers found the two model cards' headline accuracy figures came from different benchmarks and were not comparable

Laya model card ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

David Klotz (IAAI, Media University Stuttgart) April 29, 2026 (arXiv:2604.26482v1) data

Mission-critical systems of record: Retain Buy as the primary option. Consider Make selectively for peripheral modules, extensions, or integration layers where the core system's integrity is not at risk.

Agentic AI shifts make-vs-buy by application type: commodity and differentiating apps move toward build, while regulated and mission-critical systems stay buy.

3 more excerpts
  • Commodity utilities: Default to Make. Evaluate Buy only where ecosystem integrations provide strong network value or where the firm's AI capability is below the viability threshold.
  • Where software development once required large teams working over months, small teams augmented by AI agents can now deliver functional applications in days or weeks.
  • AI-era Make demands skills in prompt engineering, agent orchestration, AI output validation, and governance of AI-generated artifacts.

The Buy-or-Build Decision, Revisited: How Agentic AI Changes the Economics of Enterprise Software ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision

David Wood (O'Reilly) August 2009 (book publication) data

Fully 60% of the life cycle costs of software systems come from maintenance, with a relatively measly 40% coming from development.

Maintenance dominates software lifecycle cost, and most of that maintenance is new enhancement work rather than bug-fixing.

1 more excerpt
  • During maintenance, 60% of the costs on average relate to user-generated enhancements (changing requirements), 23% to migration activities, and 17% to bug fixes.

The 60/60 Rule (ch. 34, 97 Things Every Project Manager Should Know) ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision

DORA (Google Cloud) 2024 (page last updated April 13, 2026) data partial

AI adoption significantly increases individual productivity, flow, and job satisfaction. However, it also negatively impacts software delivery stability and throughput

AI helps the individual developer but hurts system-level delivery stability and throughput.

1 more excerpt
  • Unstable organizational priorities cause meaningful decreases in productivity and substantial increases in burnout.

Accelerate State of DevOps Report 2024 ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision

GitClear January 2026 (research notation on page) data

the percentage of changed code lines (associated with refactoring) sunk from 25% of changed lines in 2021, to less than 10% in 2024, while lines classified as 'copy/pasted' (cloned) rose from 8.3% to 12.3%

AI-assisted development correlates with more code duplication and less refactoring, increasing long-term maintenance burden on code you own.

3 more excerpts
  • 211 million changed lines from repos owned by Google, Microsoft, Meta, and enterprise C-Corps
  • 4x more code cloning
  • 'copy/paste' exceeds 'moved' code for first time in history

AI Copilot Code Quality: 2025 Data Suggests 4x Growth in Code Clones ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision

Giuseppe Destefanis, Tomaso Aste Aug 17, 2026 data

Across 1,902 instrumented runs plus a 244-run sealed-environment replication, naming one agent as coordinator in its prompt creates no communication hub and no reliable improvement in success; shared files cut output tokens by about 42% at eight agents on message-heavy work.

The cheapest fix for multi-agent coordination failures, a role label in one agent's prompt, does not work; structural changes (shared artifacts) do.

2 more excerpts
  • The null result on coordinator-naming reproduced in the sealed 244-run replication, not just the primary 1,902-run set
  • Scope is a coding-agent benchmark with a fixed test suite and specific team configurations; the authors do not claim broader generality

When Agents Coordinate: Measuring Coordination in Multi-Agent AI Coding ↗·Cited in Coordination Is an Architecture Layer, Not a Prompt Instruction

Hoagy Cunningham, Alwin Peng, Jerry Wei, Euan Ong, Fabien Roger, Linda Petrini, Misha Wagner, Vladimir Mikulik, Mrinank Sharma 2025 data

Using Claude 3.5 Haiku as a safety filter for Claude 3.5 Sonnet increases inference costs by approximately 25%.

The tier a decision model displaces already has a published price: if a small LLM is doing your filtering today, that is roughly a quarter of your inference bill, which is the baseline a vendor's speed multiple is quietly measured against.

4 more excerpts
  • EMA linear probes outperform, at negligible cost, a dedicated classifier with 2% of the policy model's parameters
  • Two-stage pipelines reduce cost by over 10x without significantly reducing overall system performance
  • The page states performance at low false-positive rate is the production-relevant metric, because that determines suitability for deployment
  • The page carries no publication date; only the 2025 path segment in the URL

Cost-Effective Constitutional Classifiers via Representation Re-use ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Hoagy Cunningham, Jerry Wei, Zihan Wang, Andrew Persic, Alwin Peng, Jordan Abderrachid, Raj Agarwal, Bobby Chen, Austin Cohen, Andy Dau, Alek Dimitriev, Rob Gilson, Logan Howard, Yijin Hua, Jared Kaplan, Jan Leike, Mu Lin, Christopher Liu, Vladimir Mikulik, Rohit Mittapalli, Clare O'Hara, Jin Pan, Nikhil Saxena, Alex Silverstein, Yue Song, Xunjie Yu, Giulio Zhou, Ethan Perez, Mrinank Sharma January 8, 2026 data

Since exchanges flagged by the first stage are escalated rather than refused, the first-stage classifier can flag a higher proportion of production traffic without incurring an excessive refusal rate. The first-layer probe escalated approximately 5.5% of traffic to the second-stage classifier, for approximately a 40x reduction compared to the single exchange classifier.

Escalating rather than refusing changes which false-positive rate you can tolerate, which is why the band you choose matters more than the accuracy number the vendor quotes.

4 more excerpts
  • The refusal flag rate was 0.05%, down from the 0.38% reported for the prior system
  • 1,736 cumulative hours of red-teaming effort across approximately 198K attempts
  • Both systems are trained on synthetic data using a constitution related to CBRN weapons, so the domain is narrower than general-purpose triage
  • An earlier research pass recorded the escalation rate as roughly 10 percent; the paper states approximately 5.5 percent

Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Kanghee Park, Jiayu Wang, Taylor Berg-Kirkpatrick, Nadia Polikarpova, Loris D'Antoni 2024 data

Constrained decoding techniques can distort the LLM's distribution, leading to outputs that are grammatical but appear with likelihoods that are not proportional to the ones given by the LLM, and so ultimately are low-quality.

Guaranteeing that an output conforms to a schema is not the same as guaranteeing the output is right; constraining the answer space can actively degrade which valid answer you get.

1 more excerpt
  • The authors propose grammar-aligned decoding (ASAp) to preserve grammaticality while matching the model's conditional distribution, so the distortion is a known and addressable artifact rather than an inherent cost

Grammar-Aligned Decoding ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Kilian Hendrickx, Lorenzo Perini, Dries Van der Plas, Wannes Meert, Jesse Davis July 23, 2021 data

Machine learning models always make a prediction, even when it is likely to be inaccurate. This machine learning sub-field was already studied in 1970 by Chow and Hellman.

Banding a model's probability with an operator-chosen threshold is the textbook reject option with a fifty-year literature, not a new capability introduced by a decision model.

3 more excerpts
  • The canonical formalism is a single threshold with two outcomes, so a three-band act/review/escalate structure is an extension of it rather than the base model
  • The standard cost ordering Cc < Cr < Ce carries a second constraint frequently dropped in summary: Cr must be no greater than 1/K for K classes
  • Ambiguity rejection and novelty rejection need different confidence functions: a class-posterior score signals ambiguity, while novelty needs a density or distance measure against the training distribution

Machine Learning with a Reject Option: A survey ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Maksym Nechepurenko, Pavel Shuvalov May 5, 2026 data

Naming and isolating the coordination layer moves four classes of properties from implicit-in-code to explicit-in-spec (failure-mode signatures, cross-system comparability, agent heterogeneity), and the specification can be implemented atop AutoGen, CrewAI, LangGraph, AWS Strands, or Microsoft Foundry without modification.

Coordination logic should be treated as its own configurable, separable architectural layer rather than embedded in agent prompts, because separation is what makes the layer analyzable at all.

3 more excerpts
  • The paper's own experiment (five coordination configurations, 100 post-cutoff Polymarket markets, Murphy decomposition) tests forecasting calibration signatures, not production failure rates
  • The commonly-cited 41-87% production-failure-rate figure appears in this paper's introduction as a citation to Cemri et al. 2025, not as this paper's own finding
  • Pairwise comparisons between coordination configurations do not survive Bonferroni correction at n=100; the authors call this a methodology-validating first instantiation, not a general cross-model claim

Coordination as an Architectural Layer for LLM-Based Multi-Agent Systems ↗·Cited in Coordination Is an Architecture Layer, Not a Prompt Instruction

Mrinank Sharma, Meg Tong, Jesse Mu, Jerry Wei, Jorrit Kruthoff, Scott Goodfriend, Euan Ong, Alwin Peng, Raj Agarwal, Cem Anil, Amanda Askell, Nathan Bailey, Joe Benton, Emma Bluemke, Samuel R. Bowman, Eric Christiansen, Hoagy Cunningham, Andy Dau, Anjali Gopal, Rob Gilson, Logan Graham, Logan Howard, Nimit Kalra, Taesung Lee, Kevin Lin, Peter Lofgren, Francesco Mosconi, Clare O'Hara, Catherine Olsson, Linda Petrini, Samir Rajani, Nikhil Saxena, Alex Silverstein, Tanya Singh, Theodore Sumers, Leonard Tang, Kevin K. Troy, Constantin Weisser, Ruiqi Zhong, Giulio Zhou, Jan Leike, Jared Kaplan, Ethan Perez January 31, 2025 data

These classifiers also maintain deployment viability, with an absolute 0.38% increase in production-traffic refusals and a 23.7% inference overhead.

A synthetic-data-trained classifier was already on the production hot path eighteen months before Jev launched, and its owners published the tax rather than only the accuracy.

4 more excerpts
  • The classifiers were trained on synthetic data generated by prompting LLMs with natural language rules, i.e. a constitution, which is the same training lineage a synthetic-only decision model sits in
  • Over 3,000 estimated hours of red teaming across 405 HackerOne participants, with bounties up to $15K per jailbreak report
  • The prototype's prioritization of robustness led to impractically high refusal rates, which is the tradeoff the next generation was built to fix
  • This sits at the input/output safety-filter layer rather than a general-purpose routing tier, so the analogy is about tax disclosure precedent rather than an identical architectural slot

Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Rachel Stephens (RedMonk) November 26, 2024 data

if AI adoption increases by 25%, estimated throughput delivery is expected to decrease by 1.5%

Individual AI productivity gains do not translate into system-level delivery throughput or stability, because code generation was never the bottleneck.

3 more excerpts
  • estimated delivery stability is expected to decrease by 7.2%
  • 75.9% of respondents (of roughly 3,000 people surveyed) are relying on AI for at least part of their job responsibilities
  • if AI adoption increases by 25%, time spent doing valuable work is estimated to decrease 2.6%

DORA Report 2024 – A Look at Throughput and Stability ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision

Sahana Chennabasappa, Cyrus Nikolaidis, Daniel Song, David Molnar, et al. May 6, 2025 data

Approximately 90% of inputs are fully resolved by the first layer, maintaining a typical end-to-end latency of under 70 milliseconds; for the remaining 10% requiring deeper inspection, end-to-end latency can exceed 300 milliseconds. PromptGuard 2 runs at 19.3ms (22M) and 92.4ms (86M).

A production agent stack already layers a deterministic tier, a small-classifier tier, and an LLM auditor, each with a measured latency, which is the real comparison set for a new decision model rather than a frontier LLM.

4 more excerpts
  • The first tier's own scan time is approximately 60 milliseconds, which is a separate measurement from the under-70ms end-to-end figure for the 90 percent it resolves
  • Named tiers: CodeShield for static analysis over 50+ CWEs in seven languages, PromptGuard 2 as the DeBERTa-based classifier, and AlignmentCheck as a chain-of-thought LLM auditor
  • Threshold selection is reported as a fixed minimal utility reduction of 3% in the AgentDojo evaluation, which is a separate evaluation from the recall-at-1%-FPR classifier table
  • Role-scoped scanning: PromptGuard analyzes only user and tool messages while AlignmentCheck evaluates assistant messages

LlamaFirewall: An open source guardrail system for building secure AI agents ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Sheryl Estrada (Fortune) August 18, 2025, 6:54 AM ET data

Purchasing AI tools from specialized vendors and building partnerships succeed about 67% of the time, while internal builds succeed only one-third as often.

Most enterprise GenAI builds fail; buying and partnering succeeds roughly three times more often than building internally.

3 more excerpts
  • 95% failure rate for enterprise AI solutions
  • about 5% of AI pilot programs achieve rapid revenue acceleration
  • 150 interviews with leaders, a survey of 350 employees, and an analysis of 300 public AI deployments

MIT report: 95% of generative AI pilots at companies are failing ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision

Tristan Farran March 13, 2026 data

Classical calibration metrics include expected calibration error, reliability diagrams, and proper scoring rules. These both provide static calibration assessment but do not address sequential monitoring with false alarm control.

Measuring calibration once is not monitoring it, and checking it repeatedly with a naive fixed-sample test manufactures false alarms, so threshold drift needs a sequential instrument rather than a recurring spot check.

1 more excerpt
  • A practitioner checking calibration daily at p<0.05 will almost certainly observe spurious alarms over a year even if calibration remains stable, because classical hypothesis tests assume a fixed sample size determined before seeing data

When Your Model Stops Working: Anytime-Valid Calibration Monitoring ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Wittawat Jitkrittum, Neha Gupta, Aditya Krishna Menon, Harikrishna Narasimhan, Ankit Singh Rawat, Sanjiv Kumar July 6, 2023 data

Such confidence-based deferral often works remarkably well in practice, but post-hoc deferral significantly improves on it where downstream models are specialists, samples carry label noise, or there is distribution shift between the train and test set.

A confidence threshold tuned in evaluation is not guaranteed to survive production traffic drift, which is the named condition under which routing on a cheap model's confidence stops being sufficient.

2 more excerpts
  • The paper's position is that confidence deferral is usually fine with three named exceptions, not that it is generally broken
  • The Bayes-optimal rule is a difference over both models' correctness probabilities against the deferral cost, so it requires the downstream model's expected correctness, which the first model's raw confidence score alone cannot supply

When Does Confidence-Based Cascade Deferral Suffice? ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D Sculley, Sebastian Nowozin, Joshua V. Dillon, Balaji Lakshminarayanan, Jasper Snoek 2019 data partial

We find that traditional post-hoc calibration does indeed fall short, as do several other previous methods.

Calibration established on one distribution does not automatically hold when the distribution moves, which is why a vendor's calibration figure has to be re-derived on your own traffic rather than adopted.

3 more excerpts
  • Population is 2019 image, text and tabular classifiers under synthetic corruptions, with no LLMs and no synthetic-training-data regime, so applying it to a decision model is an inference and is labelled as such in the post
  • The paper's headline positive result is that ensembles stay well calibrated under the same shift, so calibration is joint over model and distribution
  • A quote carried by an earlier automated research pass could not be located in the paper and was discarded; only the abstract sentence above is used

Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Yonatan Geifman, Ran El-Yaniv May 23, 2017 data

Using our method an unprecedented 2% error in top-5 ImageNet classification can be guaranteed with probability 99.9%, and almost 60% test coverage.

Buying a guaranteed risk level costs coverage, and the price is large and measurable: roughly 40 percent of inputs go unanswered to hold a 2 percent error bound.

2 more excerpts
  • Coverage is formally the probability mass of the non-rejected region, and selective risk is defined only over the accepted region normalized by coverage
  • CIFAR-10 reaches 1% error at 78.56% coverage and CIFAR-100 18.85% error at 67% coverage, both at delta=0.001

Selective Classification for Deep Neural Networks ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Yuanbo Zhou, Changjia Zhu, Junyu Wang, Xu He, Yan Zhai, Kun Sun, Mingkui Wei, Junjie Xiong May 22, 2026 data

We identify a critical blind spot arising from the mismatch between the limited inspection windows of guardrail models and the substantially larger context inference windows of downstream LLMs.

A decision model's request budget is an attack surface rather than a quota: when the inspector's window is far smaller than the actor's, the gap itself is the vulnerability.

3 more excerpts
  • Guardrails with an effective 512-token window sit in front of models accepting up to 400k tokens, with interleave-layout bypass rates consistently above 99.5% against Prompt Guard 86M
  • The population is lightweight prompt-injection and toxicity classifiers, not calibrated multi-class decision models, so this is a reason to interrogate a decision model's budget rather than a measurement of one
  • Gemini 3 Pro recognized malicious intent in all 200 tested overflow cases, raising the question of whether the actor is a better last-line detector than an undersized upstream guard

Prompt Overflow: What the Guardrail Inspects Is Not What the Model Infers ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Zimo Ji, Zongjie Li, Wenyuan Jiang, Yudong Gao, Shuai Wang April 4, 2026 data

Of the 253 actions, 93 (36.8%) were performed via Edit or Write tool calls, which are routed through Tier 2 and are not evaluated by the classifier.

A decision gate's coverage is a distinct failure channel from its accuracy: actions that never reach the model cannot be scored by it at all, and the gate's boundary has to match the actor's real action space rather than its most common channel.

3 more excerpts
  • The authors explicitly state the 81.0% stress-test FNR is not directly comparable to Anthropic's reported 17% on organic production traffic, calling the delta workload sensitivity rather than a flaw in the vendor's measurement
  • Coverage is the minority cause, not the dominant one: 51 of 115 false negatives came from the coverage gap and 64 from the classifier misjudging actions it did evaluate
  • Scope is a single model (Sonnet 4.6), one threat category, four synthetic DevOps task families with shimmed CLIs; the paper never mentions hallucination or schema conformance

Measuring the Permission Gate: A Stress-Test Evaluation of Claude Code's Auto Mode ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

a2aproject (Linux Foundation, Google-originated) practitioner

A2A's opacity principle lets agents collaborate without sharing internal memory, proprietary logic, or specific tool implementations.

A shipping, cross-vendor-governed protocol already architects coordination as a separable layer via explicit design intent.

1 more excerpt
  • Cited as design intent only, not adoption evidence: no adopter case studies or production-deployment list found on the page, and the roadmap still lists authorization schemes, dynamic capability negotiation, and streaming reliability as unresolved

Agent2Agent (A2A) protocol ↗·Cited in Coordination Is an Architecture Layer, Not a Prompt Instruction

Anthropic Jan 23, 2026 practitioner

'Planning, implementation, and testing of the same feature share too much context' to split across agents, and 'Components requiring constant back-and-forth belong in the same agent.'

There are principled places not to cut the graph - shared context and high synchronization needs are the signals to keep work in one node.

1 more excerpt
  • Used in the post as the 'where not to cut' check in the pricing list, a counterweight to over-decomposition

Building multi-agent systems: When and how to use them ↗·Cited in Task Decomposition for AI Coding Agents: Draw the Graph First

Anthropic January 9, 2026 practitioner

Because flagged exchanges are escalated to the more powerful model, rather than refused, the first-stage classifier can afford a higher false-positive rate and not frustrate the user with refusals. Runs at just ~1% additional compute cost.

The cost of the cheap decision tier is not fixed: the same vendor repriced it from 23.7 percent overhead to roughly 1 percent in a year, so a vendor's cost comparison is a snapshot of a moving baseline.

3 more excerpts
  • Refusal rate of 0.05% on harmless queries, an 87% drop from the original classifiers system
  • The classifiers were trained on synthetic data generated from a constitution of natural language rules
  • The scoping qualifier about Claude Sonnet 4.5 traffic over one month was not isolated verbatim during verification and should be re-checked before being quoted

Next-generation Constitutional Classifiers: More efficient protection against universal jailbreaks ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Bryan Ross (GitLab, Field CTO) March 24, 2026 practitioner partial

For a team of roughly 200 developers, an internal build typically costs around $1.4M in year one, requires 2–3 dedicated FTEs to maintain, and takes 12–18 months to reach a first real use case.

Building an internal agentic AI platform in regulated industries is a multi-year, multi-FTE commitment with governance surface most organizations underestimate.

2 more excerpts
  • Every engineer building the platform is an engineer _not_ modernizing a legacy pipeline, remediating security debt, or accelerating a critical delivery program.
  • Building an internal agentic AI platform in banking or insurance is a multi-year platform engineering commitment with regulatory surface area most organizations underestimate

The real cost of build vs. buy for agentic AI in regulated industries ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision

Cloudflare 2026 practitioner

Jev is TypeSafe's structured evaluation model. It evaluates one state against typed Noul, Choice, and Score questions and returns calibrated answers with probabilities and confidence. Context window listed as 32,000 tokens.

A third host independently describes the model's contract and publishes a 32,000-token context window, corroborating the cross-host discrepancy on the request budget.

1 more excerpt
  • Cloudflare does not publish per-token rates on the model page, deferring to the dashboard

Jev (typesafe) ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Digital Applied Team July 1, 2026 practitioner partial

The 2026 build-vs-buy question is less _can we afford to build it_ and more _what happens to us if the vendor moves_.

Cheaper agentic builds plus rising SaaS lock-in and repricing risk tilt the case toward owning differentiating workflows.

3 more excerpts
  • 2,698 SaaS M&A transactions closed in 2025, up 28% year over year
  • 68% of tech leaders plan vendor consolidation in 2026
  • organizations trapped in vendor lock-in face switching costs around 16 times higher

Build vs Buy: The 2026 Case for Custom AI Tools ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision

Flavio Copes September 17, 2026 practitioner

The response reports the versioned ID that answered, so log it, and pin that ID once you've tuned thresholds against it.

Threshold drift has a version axis: a floating alias moves the model your thresholds were tuned against, and nothing in your logs records that it happened unless you pin and log the resolved ID.

4 more excerpts
  • He reads the vendor's numbers carefully, noting the docs use 0.5 as a review floor and 0.9 before a destructive action as examples rather than defaults
  • He documents hands-on accuracy limits: unreliable at counting, comparing numbers, and date or window comparisons, with indirection and large irrelevant state both degrading accuracy
  • He notes scores are not a measurement and are usable to threshold or rank rather than to interpolate magnitude
  • The origin returns HTTP 403 to direct fetches and web.archive.org was blocked in this environment; the source was recovered via a reader proxy

A deep dive into Jev, TypeSafe's System One model ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

HatchWorks (Matt Paige) January 28, 2026 practitioner

AI has dramatically reduced the cost of creating software, but it hasn't eliminated the cost of owning software.

AI lowers the cost to build software but not the ongoing cost of owning and operating it, which is where build-vs-buy now turns.

3 more excerpts
  • the last 20% (security, governance, observability, performance, reliability, data quality, change management) is still 80% of the effort
  • In 2026, most enterprises land on 'yes to both.' They buy the heavy core, build what differentiates, and use AI to accelerate the glue layer.
  • If the capability is your advantage, meaning revenue, margin, speed, or defensible differentiation (AI copilots, agentic workflows, decision support)

The Build vs Buy Framework in the Age of AI ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision

Joel Spolsky October 14, 2001 (update noted December 5, 2016) practitioner

If it's a core business function — do it yourself, no matter what.

Core, business-specific functions should be built in-house because that is where control and competitive advantage live.

2 more excerpts
  • There's no way it's going to be as flexible as what Amazon does with obidos, which they wrote themselves.
  • Pick your core business competencies and goals, and do those in house.

In Defense of Not-Invented-Here Syndrome ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision

Kate Johnson Aug 25, 2026 practitioner

A storage-agnostic handoff contract (objective and ownership, a five-state machine, observations vs. mutations, verification evidence, approval gates) needs no new service, a JSON file, a database row, or a session handoff can carry it.

A concrete, lightweight artifact is what a separated coordination layer looks like in practice, not a new orchestration platform.

1 more excerpt
  • Single-author field pattern, not peer-reviewed or independently corroborated; paired with Cemri et al.'s Task Verification failure category as the academic anchor for the verification half of the claim

A small handoff contract for multiple coding agents ↗·Cited in Coordination Is an Architecture Layer, Not a Prompt Instruction

Martin Fowler July 29, 2010 (updated April 7, 2016) practitioner

for a strategic function you don't want the same software as your competitors because that would cripple your ability to differentiate.

Strategic, differentiating software should be built while commodity utility software should be bought, and the two demand different postures.

3 more excerpts
  • The 80/20 rule applies, except it may be more like 95/5
  • This is not a static dichotomy. Business activities that are strategic can become a utility as time passes.
  • For a utility function you buy the package and adjust your business process to match the software.

Utility Vs Strategic Dichotomy ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision

Matt Crabtree September 16, 2026 practitioner

Jev averages 67.8% agreement with the reference answers, where the reference answers are themselves average predictions of GPT-6 Astra and Anthropic's Fable.

The headline accuracy figure measures agreement with a synthetic consensus of two other vendors' models rather than with ground truth, which is a weaker claim than accuracy parity.

3 more excerpts
  • The piece states all numbers come from TypeSafe's own evaluation tables with no large-scale independent reproduction, and should be treated as vendor-reported
  • The workflows were written by TypeSafe's own model capabilities team, which the author flags as a possible source of bias
  • Comparison figures cited: GPT-5.6 Terra at 67.9% accuracy and 10.1s, Claude Opus 5 at 73.1% and 37.8s, with Jev reporting 0% structured output error

Jev: TypeSafe's System One Model Explained ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

OpenAI 2026 practitioner

Treat moderation scores as signals for your application's policy, not as an automatic blocking decision. We plan to continuously upgrade the moderation endpoint's underlying model. Therefore, custom policies that rely on category_scores may need recalibration over time.

A vendor shipping this exact product shape, a per-category confidence score plus a typed flag, has been telling customers for years that the threshold is theirs and will need recalibration when the model moves.

2 more excerpts
  • The endpoint is free to use, so the precedent is not confounded by pricing pressure
  • The docs never use the words alias or floating for omni-moderation-latest; that characterization is the post's analysis, supported in substance by the recalibration warning

Moderation ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

OpenAI and ROOST 2025 practitioner

OpenAI uses small, high-recall classifiers to determine if content is domain-relevant to priority risks before evaluating that content with gpt-oss-safeguard. Traditional classifiers have lower latency and cost less to sample from.

The cheap classifier tier and the expensive reasoning tier are complements rather than substitutes, stated by a vendor that ships both.

4 more excerpts
  • Traditional classifiers trained on thousands of examples will likely perform better on a task than the reasoning model
  • Rules alone handle deterministic cases well, such as keyword matches and metadata thresholds, but struggle with satire, coded language, or nuanced policy boundaries
  • The guide advises pre-filtering content sent to the reasoning model because it is more time and compute intensive than other classifiers
  • No clean general statement exists on the page instructing escalation of ambiguity to a human; only policy-template example text

User guide for gpt-oss-safeguard ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

OpenRouter 2026 practitioner

typesafe/jev-1.13 lists input at $0.042/M and output at $0/M with 32K context; the jev-latest entry states it always redirects to the latest model in the Jev family.

An independent host confirms both the pricing and the floating-alias behaviour, and publishes a context figure that does not match TypeSafe's own documentation.

1 more excerpt
  • TypeSafe's own Models page states 64k tokens per request with 32k for state plus the longest question, while OpenRouter and Cloudflare both publish 32K, so the number a reader designs against depends on which host they read

Typesafe models ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Pat Brans (CIO.com) December 11, 2025 practitioner partial

With such a layer in place, the build-versus-buy question fragments, and CIOs might buy a vendor's persona agent, build a specialized risk-management agent, purchase the foundation model, and orchestrate everything through a platform they control.

The industry consensus has shifted to hybrid: assemble build and buy across the AI stack under an orchestration layer you control.

2 more excerpts
  • Six months ago many were experimenting, but now they're scaling.
  • including cases where a senior executive's data surfaced in a junior employee's query.

Your next big AI decision isn't build vs. buy — It's how to combine the two ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision

sean goedecke September 16, 2026 practitioner

To me, this seems like a semantic dodge, since Jev can absolutely still pick the wrong choice (e.g. calling the sky "red"). Still, all of this is also true about regular LLMs with structured outputs, and it doesn't make Jev any more reliable in practice.

The zero-hallucination claim is about schema conformance, a property any LLM with structured output already has, and it says nothing about whether the decision is correct.

4 more excerpts
  • He reports no evidence the returned probabilities are anything other than regular logit probabilities
  • His own test on Qwen2.5-1.5B-Instruct with prefix plus constrained single-token decoding produced a 2x-3x speedup, suggesting the speed is an inference-stack result rather than a new architecture
  • Not being able to use test-time compute is a real ceiling that likely caps this model class around the strength of non-reasoning LLMs
  • His position is that Jev is useful but not architecturally novel, not that it is worthless

Jev means structured output is interesting again ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

sean goedecke September 20, 2026 practitioner

For serious work, a specific hand-built classifier will always be cheaper and faster than Jev. Generic classifiers have to encode knowledge of all kinds of irrelevant things in their weights, so they can address lots of different tasks. That makes them larger, slower, and more expensive to run.

A generic decision model may earn its slot only while you are still unsure the feature works, because its own logged input and output pairs become the training set for the bespoke classifier that replaces it.

3 more excerpts
  • The post is explicitly speculative, framed as what the author expects to become a common pattern, so a finite shelf life is a prediction rather than established economics
  • The bespoke path still requires ML expertise, merely deferred until after the feature is validated, so what is removed is the dataset obstacle and not the skill obstacle
  • The post says nothing about version pinning, rate limits, API shape or ownership economics

System One models like Jev can train their own replacements ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Simon Willison 12th July 2025 practitioner

The factor that stands out most to me is that these developers were all working in repositories they have a deep understanding of already, presumably on non-trivial issues since any trivial issues are likely to have been resolved in the past.

AI's edge is smallest exactly where you own and deeply understand a mature codebase long-term.

3 more excerpts
  • 56% had never used Cursor before the study
  • Developers accepted less than 44% of AI generations
  • A quarter of the participants saw increased performance, 3/4 saw reduced performance

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision

Simon Willison 11th March 2025 practitioner

it's not about getting work done faster, it's about being able to ship projects that I wouldn't have been able to justify spending time on at all.

AI's clearest payoff is enabling marginal projects that were never worth building before, not accelerating core work.

1 more excerpt
  • I'm certain it would have taken me significantly longer without LLM assistance—to the point that I probably wouldn't have bothered to build it at all.

Here's how I use LLMs to help me write code ↗·Cited in Build vs. Buy Agentic AI: Ownership Is the New Decision

Simon Willison October 6, 2025 practitioner

I can only focus on reviewing and landing one significant change at a time, but I'm finding an increasing number of tasks that can still be fired off in parallel without adding too much cognitive overhead to my primary work.

Human review-and-land throughput — one significant change at a time — is the real ceiling on how far parallel agents scale.

1 more excerpt
  • Code that started from your own specification is a lot less effort to review.

Embracing the parallel coding agent lifestyle ↗·Cited in When One Agent Stops Being Enough: The Isolation Gate

Tim Fernholz September 18, 2026 practitioner

At the end of the day, it delegates the hallucination problem a little bit to the user. The user has to say, okay, if this only comes back with 50% probability, maybe this is a coin toss, and I disregard it. But if it's 95%, sure, then I can do something with it.

Buying a calibrated decision model transfers the judgment, rather than removing it: the vendor guarantees the shape of the answer and the operator inherits the threshold that decides what to do with it.

2 more excerpts
  • The quote is Armin Ronacher, CTO of Earendil; no first-party post on Jev exists on his own site, so this article is the primary of record
  • Diogo Almeida on the training approach: an early bet that we will be making all of our data, described as one of the best bets he has made

A new kind of AI model from a ChatGPT inventor is thrilling developers ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

TypeSafe AI 2026 practitioner

Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct.

The vendor's own documentation limits the claim its launch post and the downstream coverage make: a calibrated probability is an aggregate property, so a single Jev answer can still be wrong.

3 more excerpts
  • Defines the class: System One models are built to make fast, structured decisions software can use directly, and do not write replies, produce code, or explain their reasoning
  • Confidence is returned so the caller can decide when to act and when to escalate to a person or a reasoning model
  • The page does not define RLCD and does not state cost-scaled confidence bands, despite both being attributed to TypeSafe across the coverage

System One ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

TypeSafe AI 2026 practitioner

An alias moves when a new release ships, so the answers behind it can change without a change on your side. jev-latest resolves to jev-1.13.0; 64k tokens per request with 32k for state plus the longest question; 250,000 tokens per second and 1,200 requests per minute; input $0.042/M, output free.

Adopting a hosted decision model behind a floating alias means the thing your thresholds were tuned against can change with no deploy on your side.

3 more excerpts
  • Rate limits are adjusting dynamically and the published limits can change without notice
  • Text only at launch: string, JSON object, or array of text values, with no image, audio, or video input
  • TypeSafe publishes 64k per request while Cloudflare and OpenRouter both publish 32K context, so the hosts do not agree on the number a reader would design against

Models ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

TypeSafe AI 2026 practitioner

Test thresholds by plotting confidence against accuracy on your data.

The vendor itself instructs callers to derive thresholds locally rather than adopt a published number, which is the strongest available support for treating the threshold as operator-owned.

3 more excerpts
  • Guidance is two-outcome: make code take different actions for confident and unconfident answers, escalating uncertain cases to a person or a more expensive reasoning model
  • Code examples show a two-sided uncertainty window (0.4 < spam_risk < 0.6) and a single floor (confidence < 0.8 routes to human review)
  • The docs never state cost-of-being-wrong scaling; that framing is the post's extension and is labelled as such

How to build with System One ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

TypeSafe AI 2026 practitioner

POST https://api.typesafe.ai/v1/systemone takes state plus a map of named typed questions and returns answers keyed by question id; the response echoes the resolved version (jev-1.13.0) even when the request said jev-latest.

The integration is not a drop-in model swap: the request and response shape diverge from the OpenAI chat-completions convention, and the resolved version is visible in the payload, which is what makes pinning auditable.

2 more excerpts
  • Noul returns a 0-1 probability; Choice returns the selected option plus a probabilities map plus confidence; Score returns a weighted value plus a legend
  • The docs never themselves draw the contrast with OpenAI's shape; that comparison is the post's observation from the two schemas

API reference ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

TypeSafe AI September 15, 2026 practitioner

The model never makes type errors. Claims 40x-200x faster, 70ms-500ms end-to-end, $0.042/MTok input with free output, and a homepage figure of 193.6x faster and 444.6x cheaper.

The launch page's absolute claim sits against the same vendor's concepts page conceding calibration does not guarantee an individual answer is correct, so the zero-hallucination framing is about schema conformance, not correctness.

2 more excerpts
  • Benchmark provenance is vendor-run against a vendor-selected comparison: numbers come from Workflow evals testing how they compare to the average of the smartest models, in this case Astra and Fable
  • The page does not state training exclusively on synthetic data and does not define RLCD, despite both being widely attributed to it

Introducing System One Models and Jev ↗·Cited in Where a Decision Model Belongs: Placing Jev in Your Stack

Walden Yan (Cognition) June 12, 2025 practitioner

Two named principles: 'Share context, and share full agent traces, not just individual messages' and 'Actions carry implicit decisions, and conflicting decisions carry bad results.' Illustrated with a Flappy Bird clone split across two subagents whose outputs could not be reconciled.

A written handoff specification cannot capture the implicit decisions an agent makes mid-task, which is why parallel subagents without shared context produce conflicting assumptions that were never prescribed upfront.

4 more excerpts
  • 'Actions carry implicit decisions, and conflicting decisions carry bad results.' On the parallel-subagent failure: 'The actions subagent 1 took and the actions subagent 2 took were based on conflicting assumptions not prescribed upfront.'
  • The essay's actual position is broader than an argument about edge specification: 'in 2025, running multiple agents in collaboration only results in fragile systems.' It is the steel-manned opposition to any multi-agent topology, not a supporting voice.
  • Its proposed alternative is a single-threaded linear agent.
  • Cognition's April 2026 follow-up states that multi-agent systems work best when writes stay single-threaded and additional agents contribute intelligence rather than actions

Don't Build Multi-Agents ↗·Cited in The Topology You Can Review Is the Topology You Can Run, Task Decomposition for AI Coding Agents: Draw the Graph First, When One Agent Stops Being Enough: The Isolation Gate

Walden Yan (Cognition) 04.22.26 practitioner

multi-agent systems work best today when writes stay single-threaded and the additional agents contribute intelligence rather than actions

Only split once writes can stay single-threaded and added agents are read-only intelligence — parallel-writer swarms still fail.

3 more excerpts
  • most multi-agent setups in the world are limited to 'readonly' subagents
  • The practical shape is map-reduce-and-manage: a manager splits work, children execute, the manager synthesizes
  • an average of 2 bugs per PR, of which roughly 58% are severe

Multi-Agents: What's Actually Working ↗·Cited in When One Agent Stops Being Enough: The Isolation Gate

no sources match your filters