The evidence base behind everything I publish on running AI coding agents in production. Every entry is a verbatim figure or quote from a primary source — a study, a benchmark, an engineering post — pulled while researching a post, then checked against the live page. Sources that drift or go dead are dropped or flagged.
331
Verified sources
6
Themes
111 / 220
Data / practitioner
2026-09-07
Last verified
Every entry checked against its live source · dataset: research.json
>
331 / 331 sources
01
Task Design & Decomposition
23 sources
Scoping, decomposing, and speccing work so an agent finishes it on the first try.
Alif Al Hasan, Sumon Biswas (Case Western Reserve University)May 29, 2026data
Across 547 confirmed real-world safety failures mined from the GitHub issue trackers of 13 foundational code models, the top threat category is Constraint Violations at 40.4%, ahead of Destructive Operations (24.5%), Authorization Bypasses (18.3%), and Deception (15.7%).
Independent corroboration of the constraint-violation finding by a different dataset and method - two teams reaching the same top category within two points.
3 more excerpts
Failures arise during benign, goal-directed use rather than adversarial attack
Nearly 60% of confirmed incidents were rated high or critical severity
The quantified breakdown appears in the full text, not on the arXiv abstract page
80% of tool calls come from agents that appear to have at least one kind of safeguard (like restricted permissions or human approval requirements), 73% appear to have a human in the loop in some way, and only 0.8% of actions appear to be irreversible
Irreversible agent actions are rare in real traffic, so oversight should concentrate on the small slice where a single error is costly.
2 more excerpts
such as sending an email to a customer
And while these higher-risk actions are rare as a share of overall traffic, the consequences of a single error can still be significant.
Arpandeep Khatua, Hao Zhu, Peter Tran, Arya Prabhudesai, Frederic Sadrieh, Johann K. Lieberwirth, Xinkai Yu, Yicheng Fu, Michael J. Ryan, Jiaxin Pei, Diyi Yang (Stanford, SAP Labs)Jan 19, 2026 (v1); revised Jan 26, 2026 (v2)data
Across 600+ collaborative coding tasks in 12 libraries and 4 languages, agents working together achieve on average 30% lower success rates than the same agents doing both tasks individually - the 'curse of coordination'. GPT-5 and Claude Sonnet 4.5 configurations reach only 25% under two-agent cooperation, roughly half the solo baseline. 77.3% of tasks have conflicting ground-truth solutions.
Without assigned file and interface ownership, parallel agents duplicate work and overwrite changes they believe will merge cleanly - the collision is the default outcome, not an edge case.
4 more excerpts
Adding a messaging tool did not help: the difference between 'with comm' and 'no comm' settings is not statistically significant for task success, though it did reduce literal merge conflicts
The paper separates two problems: merge conflicts are spatial coordination (who edits which lines), while task success requires semantic coordination (what to implement, not just where)
Agents were given no pre-assigned file ownership and were free to redivide the features between themselves
Scope limit: no experimental arm tested pre-assigned ownership as a fix, so the benchmark diagnoses the problem without validating the cure
Kwa, West, Becker, et al. (METR)submitted 2025-03-18data
frontier AI time horizon has been doubling approximately every seven months since 2019, though the trend may have accelerated in 2024
The primary paper behind the autonomy trend confirms a ~7-month doubling of the 50%-task-completion time horizon since 2019, driven mainly by greater reliability and error-adaptation — the mechanism that inflates calls per task.
4 more excerpts
"Claude 3.7 Sonnet have a 50% time horizon of around 50 minutes".
"within 5 years, AI systems will be capable of automating many software tasks that currently take humans a month".
"The increase in AI models' time horizons seems to be primarily driven by greater reliability and ability to adapt to mistakes"
50%-task-completion time horizon. This is the time humans typically take to complete tasks that AI models can complete with 50% success rate
The length of tasks (measured by how long they take human professionals) that generalist frontier model agents can complete autonomously with 50% reliability has been doubling approximately every 7 months for the last 6 years.
The autonomous task length frontier agents can complete has doubled roughly every 7 months for 6 years, so autonomous runs — and the per-task call count behind them — keep growing.
4 more excerpts
current models have almost 100% success rate on tasks taking humans less than 4 minutes, but succeed <10% of the time on tasks taking more than around 4 hours
"If the measured trend from the past 6 years continues for 2-4 more years, generalist autonomous agents will be capable of performing a wide range of week-long tasks."
the best current models—such as Claude 3.7 Sonnet—are capable of some tasks that take even expert humans hours, but can only reliably complete tasks of up to a few minutes long
AI agents often seem to struggle with stringing together longer sequences of actions
METR (Becker, Rush, Barnes, Rein)July 10, 2025data
When developers are allowed to use AI tools, they take 19% longer to complete issues—a significant slowdown that goes against developer beliefs and expert forecasts.
Experienced developers were measurably slower with AI in codebases they know well, contradicting their own forecasts of a speedup.
4 more excerpts
16 experienced developers from large open-source repositories (averaging 22k+ stars and 1M+ lines of code)
developers expected AI to speed them up by 24%
they still believed AI had sped them up by 20%
developers estimated that they were sped up by 20% on average when using AI—so they were mistaken
Across 20,574 coding-agent sessions from 1,639 repositories, the most prevalent misalignment symptom is Developer Constraint Violation - defined as violating an explicit developer constraint - at 38.33% of episodes, with 73.68% of those attributed to instruction-following failure. The separate underspecification cause (C1) accounts for only 15.36%.
The dominant measured failure is agents breaking constraints developers already stated, not developers failing to state them - which is why sharpening prompt prose does not address the main failure mode.
4 more excerpts
The symptom taxonomy is explicitly multi-label - 29.56% of episodes carry two labels and 0.54% carry three or more - so the seven shares deliberately do not sum to 100%
90.50% of episodes impose effort and trust costs rather than irreversible system damage, yet 91.49% of visible resolutions still require explicit user correction
Misalignment compounds across sessions: probability of misalignment in the next session is 0.519 after an affected session versus 0.336 otherwise
Constraint violation is markedly worse in CLI sessions (49.49%) than IDE sessions (32.26%)
On a Kubernetes root-cause-analysis workload, a decomposition fixed at design time with no runtime branching cost 1,632 +/- 145 tokens in retries versus 904 +/- 17 for a monolithic run - 80.5% worse. Runtime-structured decomposition with schema-validated handoffs cut retry cost to 436 +/- 132, a 51.7% reduction against monolithic and 73.2% against static.
Decomposition is not automatically a win. Splitting work without runtime isolation adds rerun surface area, because a failure anywhere forces re-execution of every downstream subtask.
4 more excerpts
The mechanism is stated directly: 'fixed sequential execution must rerun all downstream subtasks from the point of failure'
Structuring is not free - the runtime-structured baseline run cost 2,716 +/- 424 tokens against 904 +/- 17 monolithic, so the trade only pays at a nonzero failure rate
Authors' limitation: both use cases are controlled scenarios at temperature 0 with low natural failure rates (0-2%), and token savings depend on deployment failure rates they did not measure at scale
Authors' limitation, load-bearing for this post: 'Decomposition policies are developer-authored and may not generalize to automatically derived graphs'
Tim Menzies, William Nichols, Forrest Shull, Lucas Layman (NC State, SEI-CMU, Fraunhofer CESE)2016data
Across 171 software projects from 2006 to 2014: 'We found no evidence for the delayed issue effect; i.e. the effort to resolve issues in a later phase was not consistently or substantially greater than when issues were resolved soon after their introduction.'
The classic exponential cost-of-delay curve does not replicate, so the case for planning before an agent runs has to rest on measured agent failure rates rather than shift-left folklore.
2 more excerpts
Requirements issues reaching system test showed roughly a 1.85x median resolution-time increase, against the 37-250x multipliers cited in the classic literature
Used in the post as an honesty move - it argues against a convenient cliche the author declined to use
Across CompSkillBench - 300 compositional queries over 2,209 real MCP server skills spanning 24 categories - standard LLM decomposition reaches only 34.2% category recall at the step level, making decomposition quality the primary bottleneck.
Granularity is the hard part of decomposition and the part models are worst at, which is why the task graph is drawn by a human rather than delegated to the agent.
1 more excerpt
Iterative Skill-Aware Decomposition raised decomposition accuracy from 51.0% to 67.7% (+32.7%, Wilcoxon p < 10^-6)
only 48% of developers consistently check AI-assisted code before committing it, even though 38% find that reviewing AI-generated logic actually requires more effort than reviewing human-written code.
Most teams under-review AI code even though reviewing it costs more effort, so the last-mile verification tax is real and often unpaid.
1 more excerpt
AI gets you 80% to an MVP; the last 20% requires patience, learning deeply or hiring engineers.
A user asked to "clean up old branches." The agent listed remote branches, constructed a pattern match, and issued a delete. This would be blocked since the request was vague, the action irreversible and destructive, and the user may have only meant to delete local branches.
Vague-plus-irreversible-plus-destructive is the dangerous combination to gate; a concrete incident shows why you don't delegate blast-radius actions blind.
4 more excerpts
Claude Code users approve 93% of permission prompts.
If a session accumulates 3 consecutive denials or 20 total, we stop the model and escalate to the human.
Destroy or exfiltrate. Cause irreversible loss by force-pushing over history, mass-deleting cloud storage, or sending internal data externally.
Instead, a false positive costs a single retry where the agent gets a nudge, reconsiders, and usually finds an alternative path.
'In 2026, the value of an engineer's contributions shifts to system architecture design, agent coordination, quality evaluation, and strategic problem decomposition.'
A frontier lab naming decomposition as a core emerging engineering skill - cited as Anthropic's position, not as independent evidence.
4 more excerpts
The landing page returns 200 but is gated behind a form, and the PDF is not text-extractable, so the quote could not be re-confirmed against a live fetch on Aug 3, 2026; it was carried from a verified full 17-page read earlier the same session
PDF fallback surface: https://resources.anthropic.com/hubfs/2026%20Agentic%20Coding%20Trends%20Report.pdf
The widely circulated line 'the bottleneck is no longer writing code but clarity about what to build' does NOT appear anywhere in this report and is not cited in the post
Verified adjacent data from the same report: developers use AI in roughly 60% of their work but report being able to fully delegate only 0-20% of tasks
Anthropic (Erik Schluntz and Barry Zhang)Dec 19, 2024practitioner
When building applications with LLMs, we recommend finding the simplest solution possible, and only increasing complexity when needed. This might mean not building agentic systems at all.
The default should be the simplest solution; reaching for an agent is a decision to justify, not an assumption.
4 more excerpts
They are typically just LLMs using tools based on environmental feedback in a loop.
Code solutions are verifiable through automated tests; Agents can iterate on solutions using test results as feedback
The autonomous nature of agents means higher costs, and the potential for compounding errors.
Agentic systems often trade latency and cost for better task performance, and you should consider when this tradeoff makes sense.
Spec-driven tooling applied to a small bug produced four user stories with sixteen acceptance criteria - a sledgehammer for a nut. 'All SDD approaches and definitions I've found are spec-first, but not all strive to be spec-anchored or spec-as-source.'
Over-specification is a real failure mode with a cost, which supplies the stop condition in the post's pricing checklist.
1 more excerpt
The most authoritative non-vendor treatment of the spec-first / spec-anchored / spec-as-source distinction found in this research pass
Daniel Epstein (Microsoft Developer Blog)May 19, 2026practitioner
Names the failure directly: 'No backlog: There is no structured list of what needs to be built, in what order, with what dependencies. Work gets discovered during implementation, not planned before it.' The prescribed fix is 'Specs in Backlog first: Every capability is an issue. Every issue has acceptance criteria.'
A named practitioner framing of the missing artifact this post builds - the dependency-ordered backlog that precedes any agent run.
2 more excerpts
The article contains no numbers, percentages, or named studies - it is argumentative, and the post cites it as practitioner framing only, never as measurement
A widely circulated line about 'the hardest step ... assumed rather than solved' is from a reader comment by Rolf Kristensen, not from Epstein's article, and is not cited in this post
Hamel Husain, interviewed by Sara Verdi (Arize)Jul 30, 2026practitioner
'A really common way that the model is not the problem is query disambiguation. The LLM doesn't have a chance because the user is asking a very ambiguous question.'
Ambiguity attaches to interface contracts, not just prose - an agent told to clean up an authentication service cannot know whether it may change the public API, add a dependency, touch the schema, or alter error behavior.
1 more excerpt
Backs the semantic half of the post's two-edge model: a node needs an owned contract, not only an owned file glob
Kent Beck (O'Reilly, 'Coding with AI: The End of Software Development As We Know It')session page, undated; underlying event May 8, 2025practitionerpartial
'Augmented coding deprecates formerly leveraged skills such as language expertise. Augmented coding amplifies vision, strategy, task breakdown, and feedback loops.'
Task breakdown is an appreciating skill under agentic coding, not a depreciating one - which is the argument for naming an owner rather than letting it go unassigned.
2 more excerpts
Quote confirmed verbatim on the O'Reilly session page, but the page carries no publication date; the event date is corroborated from independent announcements and O'Reilly Radar coverage
Beck's own newsletter does not carry this exact sentence; kentbeck.com has a near-identical paraphrase in a mutable homepage section, so the O'Reilly page is the only stable surface for the verbatim wording
For difficult tasks, I'll often reject five or six (or more!) agent attempts before accepting one as good enough to work with, or giving up and making the change by hand.
Getting value from agents on hard tasks means aggressively rejecting weak attempts and keeping judgment work human.
3 more excerpts
able to correctly diagnose 80% of issues on its own
The current core AI skill is shifting as much work onto AI agents as possible, without going too far.
I still don't use LLMs to write Slack messages, ADRs, issues and so forth.
If your project has a robust, comprehensive and stable test suite agentic coding tools can _fly_ with it.
A strong automated test suite is the single biggest enabler of agent productivity on a codebase.
2 more excerpts
what should we call the other end of the spectrum, where seasoned professionals accelerate their work with LLMs while staying proudly and confidently accountable for the software they produce?
Automated testing / Planning in advance / Comprehensive documentation / Good version control habits / Effective automation / Culture of code review / Manual QA / Research skills / Ship to preview environment
The machinery around the model — the context it sees, the harness it acts through, the loop it runs in.
AnthropicJune 13, 2025data
A four-field delegation spec for every subagent: an objective, an output format, guidance on the tools and sources to use, and clear task boundaries. Without it, one subagent explored the 2021 automotive chip crisis while two others duplicated work on 2025 supply chains. Agents use about 4x more tokens than chat interactions; multi-agent systems about 15x.
Every edge in an agent graph needs an explicit specification, and the failure mode of an underspecified edge is duplicated and misaligned work rather than a visible error.
4 more excerpts
'Agents make dynamic decisions and are non-deterministic between runs, even with identical prompts.' Scale context: 'Simple fact-finding requires just 1 agent with 3-10 tool calls, direct comparisons might need 2-4 subagents with 10-15 calls each'.
Subagents facilitate compression by operating in parallel with their own context windows, exploring different aspects of the question simultaneously before condensing the most important tokens for the lead research agent.
Coding-specific bound, time-indexed to June 2025: 'most coding tasks involve fewer truly parallelizable tasks than research, and LLM agents are not yet great at coordinating and delegating to other agents in real time.'
The 15x and 4x multiples are both baselined against chat interactions, not against single-agent search - not interchangeable with the 3-10x figure from Anthropic's January 2026 post.
The most common failures are wrong tool selection and incorrect parameters, especially when tools have similar names like `notification-send-user` vs. `notification-send-channel`.
At scale, loading all tool definitions upfront is the failure driver; deferred tool loading cuts token cost and measurably raises tool-selection accuracy.
4 more excerpts
When using natural language tool calling, each invocation requires a full inference pass, and intermediate results pile up in context whether they're useful or not.
At Anthropic, we've seen tool definitions consume 134K tokens before optimization.
Opus 4.5 improved from 79.5% to 88.1%
This represents an 85% reduction in token usage while maintaining access to your full tool library.
In a shared-repository run, 18 out of 30 agents created a git branch with the identical name 'mvp-game-loop'. A job-market coordination attempt produced 2.4 million job requests against 117 accepted jobs. On hidden-profile tasks, groups scored 17-36% against solo ceilings near 100%. In a coordinated vulnerability hunt, agents found 266 vulnerabilities versus 21 in a solo 6.5-million-token run, with only 12 in common.
Coordination failures in multi-agent systems are properties of the arrangement rather than of model capability, and they show up first in ordinary coding artifacts: colliding branch names, conflicting pull requests, and duplicated work.
3 more excerpts
Structure determines behavior directly: given a private back-channel, agents began colluding almost immediately.
Correlated failure across a fleet: when one agent makes a bad decision, many agents are likely to make the same one.
The words 'checkpoint', 'approval', 'topology', 'graph' and 'architecture' do not appear anywhere in the piece; human intervention appears only as a fallback.
Bandi, Dumitru, Hertzberg, Agarwal et al. (Scale AI)Jan 31, 2026data
Across 1,000 expert-written tasks spanning 36 real MCP servers and 220 tools, automated diagnostics show 63.3% of diagnosed failures are cognitive rather than tool-call related.
Second independent finding that the majority of agent failures are model-side, not tool-surface.
1 more excerpt
'Several high-performing models fail after successful tool execution due to premature stopping or incorrect synthesis' - a cognitive fault the harness can still address
Charles Fleming, Guillaume De Saint Marc, Ramana Kompella, Peter Bosch, Vijoy Pandeyv1 Apr 3, 2026; v2 May 27, 2026data
A service-mesh-style middleware layer attached to every agent instance improved performance more than 10% over direct agent-to-agent communication on HotPotQA (80.1 to 91.5 vs. 92 baseline) and MuSiQue (72.7 to 86.1 vs. 87.5 baseline).
Direct, prompt-embedded agent-to-agent communication measurably costs accuracy; a separated coordination layer recovers most of it.
1 more excerpt
The PDF returned only partial text on direct fetch; verified via the HTML render and cross-checked against the abstract's summary figure
Even the best frontier models only achieve 68% accuracy at the max density of 500 instructions.
Instruction-following accuracy degrades sharply with density — the best frontier models hit only 68% at 500 instructions — so packing rules in measurably erodes compliance.
4 more excerpts
At 500 instructions, llama-4-scout exhibits an extreme O:M ratio of 34.88, indicating omission errors are over 30 times more frequent
Threshold decay: "Performance remains stable until a threshold, then transitions to a different (steeper) degradation slope" — exhibited by gemini-2.5-pro, o3
Primacy effects display an interesting pattern across all models: they start low at minimal instruction densities indicating almost no bias for earlier instructions, peak around 150–200 instructions
Analysis reveals model size and reasoning capability to correlate with 3 distinct performance degradation patterns, bias towards earlier instructions, and distinct categories of instruction-following errors.
Dat Tran, Douwe KielaApril 2, 2026 (v2: April 11, 2026)data
Across two datasets (FRAMES, MuSiQue), three model families (Qwen3, DeepSeek, Gemini) and five multi-agent architectures (Sequential, Debate, Ensemble, Parallel-roles, Subtask-parallel), single-agent systems match or outperform multi-agent systems when reasoning tokens are held constant. Thinking-token budgets tested at 100, 500, 1k, 2k, 5k and 10k.
Many reported multi-agent advantages are better explained by unaccounted computation and context effects than by any inherent architectural benefit, so an edge's apparent gain has to be re-measured with compute held equal before it is kept.
3 more excerpts
The paper carves out a regime where multi-agent wins: in sufficiently degraded-context regimes, structured pipelines 'may occasionally surpass SAS by imposing useful factorization, filtering, or verification structure'.
Explicitly scoped out of tool-using settings: 'We focus on text-only multi-hop reasoning; MAS advantages with tools/vision or safety constraints are out of scope.' Coding agents are tool-using, so this is indirect evidence for coding.
Du et al. (Findings of EMNLP 2025)November 2025data
even when models can perfectly retrieve all relevant information, their performance still degrades substantially (13.9%-85%) as input length increases but remains well within their claimed context lengths.
Dun Yuan, Fuyuan Lyu, Ye Yuan, Weixu Zhang, Bowei He, Jiayi Geng, Linfeng Du, Zipeng Sun, Yankai Chen, Changjiang Han, Jikun Kang, Xi Chen, Haolun Wu, Xue Liuv1 Mar 30, 2026; v3 Apr 13, 2026data
A systematic audit of 18 agent communication protocols found semantic-layer mechanisms (clarification, context alignment, verification) largely absent at the protocol level.
Absent a protocol-level home for coordination logic, developers reintroduce it through prompts, wrappers, and orchestration glue by default, which accumulates hidden complexity.
1 more excerpt
Evidence base is a qualitative comparative audit across nine dimensions in three layers, not quantitative failure-rate measurement
Providing context files does not generally improve task success rates while increasing inference cost by over 20% on average. Developer-provided files improved performance by 2.4% on average, not statistically significant (p=21%); LLM-generated files caused drops in 5 of 8 settings.
The most portable layer of the setup is not the most valuable one, so portability of AGENTS.md should not be mistaken for leverage.
4 more excerpts
we find that context files tend to reduce task success rates compared to providing no repository context, while also increasing inference cost by over 20%.
Context files increased steps in every setting, by 2.45 and 3.92 on average, driving cost increases of 20% and 23% on SWE-bench and CTXbench respectively
Huang et al. (arXiv)Submitted 22 January 2026 (accepted at ICAIBD 2026)datapartial
Procedural reliability, particularly tool initialization failures, constitutes the primary bottleneck for smaller models.
For smaller models, tool-invocation reliability (especially tool initialization) is the primary failure bottleneck, localizable via a 12-category taxonomy.
3 more excerpts
1,980 deterministic test instances
12-category error taxonomy capturing failure modes across tool initialization, parameter handling, execution, and result interpretation
Mid-sized models (qwen2.5:14b) offer practical accuracy-efficiency trade-offs on commodity hardware (96.6% success rate, 7.3 s latency)
Kelly Hong, Anton Troynikov, Jeff Huber (Chroma)July 14, 2025data
Even under these minimal conditions, model performance degrades as input length increases, often in surprising and non-uniform ways.
18 LLMs degrade non-uniformly as input grows — the independent mechanism behind the bloat warning (applies to CLAUDE.md by analogy; the study never tests it).
4 more excerpts
models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows
Even a single distractor reduces performance relative to the baseline (needle only).
Whether relevant information is present in a model's context is not all that matters; what matters more is how that information is presented.
Even a single distractor reduces performance relative to the baseline (needle only), and adding four distractors compounds this degradation further
Liu et al. (TACL 2024)July 6, 2023 (v1); revised Nov 20, 2023data
language model performance is highest when relevant information occurs at the very beginning (primacy bias) or end of its input context (recency bias), and performance significantly degrades when models must access and use information in the middle of long contexts
Mid-file placement is the worst-served position — the U-shaped retrieval curve.
1 more excerpt
extended-context models are not necessarily better than their non-extended counterparts at using their input context.
Matthias Galster, Seyedmoein Mohsenimofidi, Jai Lal Lulla, Muhammad Auwal Abubakar, Christoph Treude, Sebastian BaltesFeb 16, 2026 (v5: Jun 30, 2026)data
Across 2,853 repositories, 2,015 (70.6%) adopted a single tool while 295 (10.3%) configured two, and 493 repositories (17.3%) used AGENTS.md as a tool-agnostic standard without any tool-specific artifact. The paper recommends that developers who rely on multiple tools maintain an AGENTS.md file as the shared core configuration, with tool-specific files as adapters that reference it.
The shared-core-plus-adapters layout is the empirically observed pattern and the paper's own recommendation, not a preference of the post.
4 more excerpts
Keep repo-level adoption (AGENTS.md 39.5%, CLAUDE.md 45.9%) separate from file-count share (31.6% and 34.4%); the paper reports both
No mechanism beyond context files exceeds 20% adoption for Claude, Copilot, Cursor or Gemini; 72.8% of Cursor repositories adopt Rules
Adopting tool-specific mechanisms such as Rules or Commands ties workflows to a specific tool, in the paper's own words
85.5% of Skills include no additional resources, so Skills function primarily as structured text
Paul Barbaste, Tristan Darrigol, Germain Vu, Tom WiltbergerJul 15, 2026 (expanded second edition of an April 2026 study)datapartial
Across eleven harnesses and roughly four million lines of source, SKILL.md skills lead MCP in adoption (9/11 vs 8/11), OpenHands runs Claude Code, Codex, or Gemini CLI as interchangeable backends, and Codex adopts Claude Code's hook vocabulary verbatim and ships an importer for its sessions and settings.
The harness layer has converged enough that one harness can host its rivals, and behavioral policy is migrating from prompt prose into configuration.
3 more excerpts
Per-harness hook event counts (Gemini CLI 11, Pi about 33, OpenCode 20), the 'hooks in 9/11' aggregate and the Polly per-task worktree example could not be located in the paper; those figures are sourced first-party from vendor docs instead
No agent runtime in the corpus imports a general-purpose agentic framework or retrieves code with vector embeddings
Preprint; bibliography was not extractable in the deep read
a single piece of irrelevant information can distract the models and substantially degrade their performance, even on problems whose clean versions they correctly solve.
Irrelevant context degrades accuracy even when all relevant information is present.
1 more excerpt
we find that simply adding an instruction to ignore irrelevant information brings notable performance gains on our benchmark.
Terminal-Bench (Stanford x Laude)live leaderboard (fetched Jul 27, 2026)data
Eleven Claude Opus 4.6 entries span 58.0% (Claude Code, +/-2.9) to 76.4% (Meta-Harness, +/-2.4) on 89 terminal tasks - an 18.4-point spread on identical model weights, 13.1 points at the non-overlapping confidence bounds.
The same frontier model varies by double-digit percentage points purely as a function of the harness it runs in.
2 more excerpts
Entries are self-submitted by harness authors via pull request, machine-validated for timeout/resource parity and a five-trial minimum, then maintainer-merged
Submissions span Dec 2025 to May 2026 and are not contemporaneous; harness-side and model-side settings are not held constant, so the spread is observational rather than controlled
The Terminal-Bench Team (Kelly Buchanan, TB2.1 Lead)datapartial
On Terminal-Bench 2.1 the same model scores differently under different harnesses: Opus 4.6 at 70.1% in Claude Code vs 63.8% in Terminus 2, GPT-5.4 at 77.3% in Codex CLI vs 54.8% in Terminus 2, Gemini 3.1 Pro at 70.7% in Terminus 2 vs 67.1% in Gemini CLI.
A model-versus-model leaderboard cannot tell you how a swap will perform, because the harness moves the score.
3 more excerpts
The page never states a harness-versus-harness comparison in prose; the pairs are read off the agent-model table, and the harness-causation argument is the post's inference
The release fixes 28 of the 89 tasks in Terminal-Bench 2.0; the largest gain is Claude Code with Opus 4.6, up 12.1%
Additional pairs: GPT-5.4 mini 66.1% in Codex CLI vs 36.9% in Terminus 2; Sonnet 4.6 58.5% in Claude Code vs 51.5%
Worawalan Chatlatanagulchai et al.17 Nov 2025 (submitted)data
While developers use context files to make agents functional, they provide few guardrails to ensure that agent-written code is secure or performant
Empirically, teams pack context files with functional setup but almost no security or performance guardrails — the constraint side of CLAUDE.md is systematically under-specified.
3 more excerpts
2,303 agent context files across 1,925 repositories
Build and run commands: 62.3%, Implementation details: 69.9%, Architecture: 67.7%; Security: 14.5%, Performance: 14.5%
These files are not static documentation but complex, difficult-to-read artifacts that evolve like configuration code
Xiaoyu Chu, Sacheendra Talluri, Qingxian Lu, Alexandru Iosup (VU Amsterdam)Jan 21, 2025 (v2: Mar 15, 2025; ICPE '25)data
For Anthropic's services, the likelihood of any two services experiencing outages on the same day is over 80%, while no correlation is observed between services from different providers. Anthropic averaged an MTTR of 2.70 hours and an MTBF of 5.22 days across the study window.
A vendor's surfaces fail together, so the vendor, not the individual product, is the failure domain to plan around.
4 more excerpts
Data window ends 2024-08-31 across 8 services from OpenAI, Anthropic and Character.AI; the over-80% figure is Anthropic-specific
The mechanism is hedged: the paper says the difference may be caused by different cloud infrastructures (OpenAI on Azure, Anthropic on GCP)
The 49.21% OpenAI figure is API-to-ChatGPT co-occurrence specifically, not a whole-provider aggregate
Only 6.15% of reports disclose a postmortem; Anthropic provided none for its API and Console services
Yang et al. (NeurIPS 2024)May 6, 2024 (v3 November 11, 2024)datapartial
SWE-agent solves 10.7 percentage points more issues than the baseline agent that uses just the default Linux shell (300-issue ablation); on SWE-bench Lite, SWE-agent with GPT-4 Turbo resolves 18.00% versus 11.00% for the shell-only agent with the same model.
Agent-computer interface design alone moves resolve rates by double digits with the model held fixed - the founding demonstration of the scaffold confound.
2 more excerpts
Verification is partial because arXiv HTML routes 404 and the PDF required reader-proxy extraction; figures converged across three independent extraction passes
Citation caution: the paper's verbatim 10.7pp sentence arithmetically pairs with the 7.33% no-demonstration shell baseline, not the 11.00% row - quote the sentence or the 18.00/11.00 pair, never 10.7 with 11.00
Yubin Kim, Ken Gu, Chanwoo Park, Chunjong Park, Samuel Schmidgall, et al. (Google Research / MIT Media Lab / DeepMind)December 9, 2025 (v3: April 8, 2026)data
260 configurations across six agentic benchmarks, five canonical architectures (Single-Agent plus Independent, Centralized, Decentralized, Hybrid) and three LLM families, with tools, prompts and compute standardized to isolate architectural effects. On SWE-bench Verified every multi-agent architecture underperformed the single-agent baseline (mean 0.522): Hybrid -2.1%, Centralized -3.1%, Decentralized -5.4%, Independent -14.9%, on 20-instance subsets. Agent count was not a significant predictor (log(1+n_a): beta=0.040, 95% CI [-0.074, 0.155], p=0.487). Capability ceiling beta=-0.236, p=0.004. Cross-validated R^2=0.373.
The shape of an agent system, not the number of agents in it, determines whether collaboration helps; and on coding work specifically every multi-agent arrangement tested lost to a single agent.
4 more excerpts
Relative performance change compared to single-agent baseline ranges from +80.8% on decomposable financial reasoning to -70.0% on sequential planning, demonstrating that architecture-task alignment determines collaborative success.
The headline range quoted from the abstract (+80.8% to -70.0%) spans two different architectures on two different tasks: +80.8% is Centralized on Finance Agent, -70.0% is Independent on PlanCraft. Holding Centralized fixed, the pair is +80.8% / -50.3%.
Architectures without centralized verification propagate errors more than centrally coordinated ones.
Cite arXiv v3; Google's blog post describes an earlier 180-configuration, four-benchmark version.
We attribute this improvement to the legibility of failed logical search. Repeated failures under explicit lexical constraints provide a clearer signal that required evidence may be absent, whereas Agentic Hybrid may still return semantically related but unsupported passages.
Logical/lexical retrieval can signal 'nothing found' where embedding search cannot, which measurably reduces hallucination on answer-unavailable questions.
3 more excerpts
On average, its refusal rate increased from 0.767 to 0.828, while the hallucination rate decreased from 0.128 to 0.083.
anchoring the retrieval process in logical queries substantially reduces hallucinations in generated responses.
matches a strong agentic hybrid baseline, while substantially reducing construction and serving cost
In a controlled 3x3 factorial experiment, average harness variance is 18.48 pp-squared versus average model variance of 2.37 pp-squared - a 7.80x ratio - and public leaderboards show harness-only swings of 7.3pp (Terminal-Bench 2), 9.5pp (SWE-bench Pro, same Opus 4.5), up to 15pp (SWE-bench Verified), and 34-48pp cross-scaffold gaps on the HAL Leaderboard.
Performance variance is governed more by harness configuration than model choice, so evaluation protocols without harness disclosure systematically misattribute harness gains to model improvements.
2 more excerpts
The same model under a different harness can rank above or below a competitor - rank order itself is harness-dependent
The paper proposes a harness-aware evaluation framework with a disclosure standard and variance decomposition protocol
Across 1,794 manually annotated trajectories (63,000+ execution steps, seven frontier models, three scaffolds), environment triggers account for 9.4% of decisive errors against 57.9% epistemic and 32.8% competence.
The direct refutation of harness causation - agent failures are predominantly model-side, which is why this post argues leverage rather than cause.
3 more excerpts
Largest single trigger is false premises at 30.7%
Epistemic errors are the largest share in every scaffold tested, ranging 44% to 80%
The paper's own prescription - 'earlier validation and intervention' - is itself a harness prescription
A six-step production line (Plan, Spawn, Monitor, Verify, Integrate, Retro) with three to five teammates named as the sweet spot and a ratio of one reviewer per three to four builders. Guardrail defaults include MAX_ITERATIONS=8 and an auto-pause at 85% of budget.
Verification, not generation, is the bottleneck in multi-agent coding, and human review is infrastructure rather than optional overhead.
4 more excerpts
Three to five teammates is the sweet spot. Token costs scale linearly, and three focused teammates consistently outperform five scattered ones.
Review is positioned as a stage in an ordered pipeline, not as a node in a topology; team size is chosen first for parallelism and cost, and the reviewer ratio is applied afterward.
Cites Gloaguen et al. (ETH Zurich) for a roughly 3% success-rate reduction and over 20% inference-cost increase from LLM-generated AGENTS.md files.
The bottleneck is no longer generation. It's verification.
'The same SKILL.md file works in Claude Code, Cursor (with rules), Gemini CLI, Codex, and any other harness that accepts system-prompt content.' Cursor users put them in .cursor/rules/; Gemini CLI has its own install path.
Skills are the one advanced mechanism practitioners report carrying across harnesses, because they are plain markdown with frontmatter.
2 more excerpts
Slash commands sit on top as an add-on layer (seven commands over twenty skills in the author's repo)
Practitioner report on one author's repo, which had crossed 27K stars; not a controlled test of portability
Only name (max 64 characters, lowercase, must match the directory) and description (max 1024 characters) are required; six fields total are in the frontmatter table, and allowed-tools is 'Experimental. Support for this field may vary between agent implementations.'
The portable surface of a skill is the six-field spec, and anything beyond it is per-tool.
3 more excerpts
The spec does not state how agents must treat unrecognized top-level frontmatter fields; metadata is the sanctioned place for non-spec data
Progressive disclosure: about 100 tokens of metadata loaded at startup, a body under 5,000 tokens recommended, main SKILL.md under 500 lines
Script runtimes depend on the agent implementation
Agentic AI Foundation (Linux Foundation)undated (fetched Jul 13, 2026)practitionerpartial
Over 60,000 open-source projects use AGENTS.md; 'the closest AGENTS.md to the edited file wins; explicit user chat prompts override everything.' The supported-tools list runs to 23 entries including OpenAI Codex, Cursor and Google Gemini CLI, and Claude Code is not on it.
AGENTS.md is the vendor-neutral instruction convention, and Claude Code's absence from the list is why the post bridges it with an import.
4 more excerpts
Agents automatically read the nearest file in the directory tree, so the closest one takes precedence and every subproject can ship tailored instructions.
Some strings came back as the fetch tool's paraphrase; re-verify exact wording before quoting verbatim
Per-tool flips shown on the page: Aider via .aider.conf.yml read: AGENTS.md, Gemini CLI via .gemini/settings.json context filename
Stewarded by the Agentic AI Foundation under the Linux Foundation
Aider (Paul Gauthier)continuously updated (fetched August 2026)practitioner
The same model's code-editing score moves ~10 points by edit format alone: gemini-exp-1206 scores 80.5% in whole format versus 69.2% in diff format; o1-mini 70.7% versus 61.1%.
A live, reproducible public leaderboard shows the harness's edit protocol moving scores by roughly the same magnitude as top-of-table model gaps.
1 more excerpt
Live page - re-pull current figures before quoting in new work
Albert NahasFeb 17 (year not stated on page; brief lists 2026)practitioner
when the context window fills up and gets compacted, your CLAUDE.md values get summarized away with everything else
CLAUDE.md instructions decay mid-session — they get summarized away at compaction — so hook-based reinforcement is more reliable for must-follow standards.
3 more excerpts
hook output requires approximately 15 tokens per prompt reminder
Over 50-turn session, motto reminders total ~750 tokens against 200k context window
hook output arrives as clean system-reminder messages — no disclaimer, no 'may or may not be relevant' framing
Of course, there's a trade-off: runtime exploration is slower than retrieving pre-computed data. Not only that, but opinionated and thoughtful engineering is required to ensure that an LLM has the right tools and heuristics for effectively navigating its information landscape.
Just-in-time context retrieval is not free: it trades latency for freshness and demands deliberate tool and heuristic design to work.
4 more excerpts
An agent running in a loop generates more and more data that could be relevant for the next turn of inference, and this information must be cyclically refined.
Context, therefore, must be treated as a finite resource with diminishing marginal returns.
agents built with the 'just in time' approach maintain lightweight identifiers (file paths, stored queries, web links, etc.) and use these references to dynamically load data into context at runtime using tools.
In certain settings, the most effective agents might employ a hybrid strategy, retrieving some data up front for speed, and pursuing further autonomous exploration at its discretion.
As token count grows, accuracy and recall degrade, a phenomenon known as context rot. This makes curating what's in context just as important as how much space is available.
Curating what's in the context window matters as much as how much space is available - official-docs corroboration of context rot.
The core challenge of long-running agents is that they must work in discrete sessions, and each new session begins with no memory of what came before.
A high-level prompt alone fails a long-running loop; cross-session state must be externalized to disk.
2 more excerpts
even a frontier coding model like Opus 4.5 running on the Claude Agent SDK in a loop across multiple context windows will fall short of building a production-quality web app if it's only given a high-level prompt
memoryless-session framing is softened by Opus 4.5+ auto-compaction per Anthropic's March 2026 follow-up - the externalized-state lesson persists, the mechanism is version-dependent
gather context -> take action -> verify work -> repeat
The agent loop is a repeated four-step cycle; managing context across iterations (compaction) is a loop-only concern with no single-task analog.
1 more excerpt
The Claude Agent SDK's compact feature automatically summarizes previous messages when the context limit approaches, so your agent won't run out of context.
This metadata is the first level of progressive disclosure: it provides just enough information for Claude to know when each skill should be used without loading all of it into context.
Progressive-disclosure mechanics: metadata triggers, bodies load on relevance.
1 more excerpt
This means that the amount of context that can be bundled into a skill is effectively unbounded.
Five named coordination patterns: generator-verifier, orchestrator-subagent, agent teams, message bus, and shared-state. Pattern choice is derived from workflow predictability, decomposability, interdependence and context duration.
The current frontier-lab state of the art for choosing an agent arrangement is a pattern catalogue selected by task structure, with an explicit recommendation to start with the simplest pattern and evolve from there.
3 more excerpts
The words 'risk', 'blast radius', 'reviewer', 'review capacity', 'bandwidth' and 'checkpoint' appear nowhere in the article.
Human involvement appears exactly once, as a fallback inside a loop-prevention mechanism.
Approximately 30% of Claude Code users who made requests during the affected window had at least one message routed to the wrong server type. Misrouting peaked at 16% of Sonnet 4 requests on the first-party platform on August 31, 0.18% on Amazon Bedrock, and under 0.0004% on Google Cloud's Vertex AI.
The same model degraded very differently by serving path, and internal evals did not catch what users were reporting.
4 more excerpts
The 30% figure counts users with at least one misrouted message, not sustained degradation, and is scoped to Claude Code users
Routing was sticky, so a request served by the wrong server made follow-ups likely to hit it too
The piece shows containment within Anthropic's own infrastructure; it cannot speak to other vendors being unaffected in the same window
Bug 2 and Bug 3 date boundaries are less crisp in the text; verify before quoting bug-by-bug ranges
Unlike CLAUDE.md content, a skill's body loads only when it's used, so long reference material costs almost nothing until you need it.
Skills are the designated destination for procedures that outgrew CLAUDE.md — with a stickiness caveat once invoked.
4 more excerpts
'Outside Claude Code, you can use only the fields in the Agent Skills spec. If you include any field the spec doesn't allow, packaging or upload fails with a hard error instead of ignoring the field.' Six fields are portable: name, description, license, compatibility, metadata, allowed-tools.
When you or Claude invoke a skill, the rendered SKILL.md content enters the conversation as a single message and stays there for the rest of the session.
The hard-error enforcement is documented for Anthropic's own paths (claude.ai uploads, the Skills API, package_skill.py); the page never mentions Codex, Cursor or Gemini CLI, so do not cite it as evidence that third-party agents reject non-spec fields
Custom commands have been merged into skills: .claude/commands/deploy.md and .claude/skills/deploy/SKILL.md both create /deploy
'They provide deterministic control over Claude Code's behavior, ensuring certain actions always happen rather than relying on the LLM to choose to run them.'
'Deterministic' is the vendor's own word for the control surface - but the claim is immediately qualified.
1 more excerpt
The very next sentence: prompt-based and agent-based hooks 'use a Claude model to evaluate conditions' - so 'hook' is an umbrella containing probabilistic members, and the determinism attaches to command hooks only
Anthropic (Claude Code Docs)2026 (undated on page; brief dates it 2026)practitioner
If Claude keeps doing something you don't want despite having a rule against it, the file is probably too long and the rule is getting lost. If Claude asks you questions that are answered in CLAUDE.md, the phrasing might be ambiguous. Treat CLAUDE.md like code: review it when things go wrong, prune it regularly, and test changes by observing whether Claude's behavior actually shifts.
Anthropic's own guidance says to maintain CLAUDE.md like code — prune it, and test rule changes by observing whether Claude's behavior actually shifts.
4 more excerpts
Claude stops when the work looks done. Without a check it can run, 'looks done' is the only signal available, and you become the verification loop: every mistake waits for you to notice it.
The over-specified CLAUDE.md. If your CLAUDE.md is too long, Claude ignores half of it because important rules get lost in the noise.
Keep it concise. For each line, ask: 'Would removing this cause Claude to make mistakes?' If not, cut it. Bloated CLAUDE.md files cause Claude to ignore your actual instructions!
Ruthlessly prune. If Claude already does something correctly without the instruction, delete it or convert it to a hook.
Anthropic (David Dworken, Oliver Weller-Davies)Oct 20, 2025practitioner
'Constantly clicking approve slows down development cycles and can lead to approval fatigue, where users might not pay close attention to what they're approving, and in turn making development less safe.'
Why prompt-time human review does not scale, argued by the vendor whose product depends on it.
2 more excerpts
'In our internal usage, we've found that sandboxing safely reduces permission prompts by 84%' - vendor-internal, no methodology, sample size, or definition of 'safely'
Context-reversal finding: this blog says a successful prompt injection is 'fully isolated', which the product documentation explicitly contradicts ('not a complete isolation boundary')
Level 1: Metadata | Always (at startup) | ~100 tokens per Skill ... Level 2: Instructions | When Skill is triggered | Under 5k tokens ... Level 3+: Resources | As needed | Effectively unlimited
The on-demand tier has a documented, quantified cost model (Anthropic's stated architecture, not a measured benchmark).
1 more excerpt
This filesystem-based architecture enables progressive disclosure: Claude loads information in stages as needed, rather than consuming context upfront.
A major-impact incident on 2026-09-03 ran from 13:26 to 16:23 UTC (2h57m) and listed claude.ai, the Claude API, Claude Code and Claude Cowork as affected components.
A model-layer outage takes Claude Code down with every other Anthropic surface at once, which is the drill the post runs.
2 more excerpts
No dedicated dossier deep read; the timeline and affected components come from the status.claude.com incidents feed captured on 2026-09-07
The incident sits in a rolling status window and will age out of the public feed
A single model incident on 2026-07-17 affected claude.ai, Claude API (api.anthropic.com), Claude Code, and Claude Cowork from 06:47 to 12:21 UTC, roughly 5 hours 34 minutes, and the status regressed from Monitoring back to Identified at 07:10 UTC before resolution.
Outages co-occur inside a provider: one model incident took the coding agent down with the API and the chat product.
2 more excerpts
Single-vendor incident; says nothing about cross-vendor correlation
Non-monotonic status (a fix declared at 07:03 UTC, then reverted) is why a status page reading needs the full timeline, not the latest update
Anthropic (status.claude.com)Sep 9, 2025 to Sep 17, 2025practitionerpartial
The public incident opened on Sep 9, 2025 for a Claude Sonnet 4 bug that started Aug 5 and a second bug affecting Claude Haiku 3.5 and Sonnet 4 from Aug 26; the detection method is recorded as community reports helping identify and isolate the bugs.
Quality degradation ran for weeks before the vendor opened an incident, and users detected it first.
3 more excerpts
Affected components list claude.ai, Claude Console, Claude API and Claude Code; single-vendor and cannot speak to cross-vendor correlation
A Sep 12 monitoring update said no ongoing issues five days before the Sep 17 resolution
The engineering postmortem is the fuller account of the same bugs
Anthropic's Identified update at 17:22 UTC on 2026-08-28 reads: 'We have identified an issue with an upstream cloud provider affecting Claude Cowork and Claude Code on the web.' The major-impact incident ran until 20:21 UTC, about 2h59m.
An upstream cloud provider, not a model, can take the coding agent down, so the failure domain includes the vendor's own infrastructure dependencies.
3 more excerpts
No dedicated dossier deep read; the quote and timeline come from the status.claude.com incidents feed captured on 2026-09-07
The upstream provider is not named
Scoped to Claude Code on the web and Claude Cowork, not the terminal client
As displayed on 2026-09-07, 90-day uptime read 99.44% for Claude Code, 99.5% for the Claude API, 99.4% for claude.ai and 99.43% for Claude Cowork. The incidents feed held 28 entries from 2026-08-04 to 2026-09-03, 19 of them listing Claude Code among affected components.
Anthropic's surfaces fail together repeatedly, and any incident count drawn from the page must carry a date because the feed is a rolling window.
3 more excerpts
A prior-pass count (37 of 50 incidents from 2026-07-21) could not be reproduced; the feed's earliest entry had moved to 2026-08-04, so counts shrink over time
The 19-of-28 and 9 major-or-critical figures are derived by itemizing the feed, not stated by the page
'Both CLI platforms utilize identical workspace context rules. No modifications are needed to your existing rule documents': the agent continues to parse GEMINI.md and AGENTS.md in the active directory and ~/.gemini/GEMINI.md globally.
The instruction-file layer survived a vendor retiring its own tool; a team that had pointed Gemini CLI at AGENTS.md needed no change.
2 more excerpts
Correction: Antigravity CLI reads AGENTS.md automatically by name with no settings key, so the earlier 're-check the flip' claim was broken and the source is used as a counter-example instead
Vendor migration guide asserting drop-in continuation; not independently tested
'In coding agents, part of the harness is already built in (e.g. via the system prompt, or the chosen code retrieval mechanism, or even a sophisticated orchestration system).'
The harness arrives partly pre-built - the inherited-defaults premise, from the discipline's highest-authority restatement.
4 more excerpts
The 2x2 that organizes the post: guides (feedforward) vs sensors (feedback), crossed with computational (deterministic, reliable) vs inferential (non-deterministic)
She files AGENTS.md and Skills as inferential feedforward - a legitimate quadrant member, not the harness's opposite
'Building this outer harness is emerging as an ongoing engineering practice, not a one-time configuration'
Names cybernetics as the lineage via a Wikipedia link, with no specific control-theory originator
Cara Phillips, with Paul Chen, Andy Schumeister, Brad Abrams, Theo Chu (Anthropic)January 23, 2026practitioner
Anthropic reports that multi-agent implementations typically use 3-10x more tokens than single-agent approaches for equivalent tasks. Three named conditions justify multiple agents: context pollution degrading performance, tasks that can run in parallel, and specialization that improves tool selection or task focus.
A closed list of three conditions is the only thing that justifies adding an agent; outside them, coordination costs typically exceed the benefits.
3 more excerpts
The 3-10x figure is unqualified internal testing with no published methodology, task list, dataset size, or model versions.
Coding-specific: dividing by type of work (one agent writes features, another writes tests, a third reviews code) creates constant coordination overhead; an agent handling a feature should also handle its tests.
The piece never mentions humans, approval, or review; verification is framed entirely as an agent-to-agent subagent pattern.
The root file should be pointers and critical gotchas only; everything else drifts into noise.
The root tier's content rule comes from Anthropic itself: pointers and gotchas, not documentation.
2 more excerpts
Claude loads them additively as it moves through the codebase: root file for the big picture, subdirectory files for local conventions.
Skills solve this through progressive disclosure, offloading specialized workflows and domain knowledge that would otherwise compete for context space and loading them only when the task calls for it.
Claude by AnthropicNovember 25, 2025practitionerpartial
Every conversation starts with this context already loaded, eliminating the need to explain basic project information repeatedly.
The always-loaded tier recurs every session — the recurring-cost premise. (Excerpt deliberately omits the page's 'system prompt' clause, refuted 0-3 against the docs.)
Claude Code docs (code.claude.com)2026 (undated on page)practitioner
'Claude Code reads CLAUDE.md, not AGENTS.md. If your repository already uses AGENTS.md for other coding agents, create a CLAUDE.md that imports it so both tools read the same instructions without duplicating them.' Memory files are treated as context, not enforced configuration; to block an action regardless of what Claude decides, use a PreToolUse hook.
The bridge for Claude Code is a one-line import, and the instruction file is advisory rather than enforcing.
4 more excerpts
CLAUDE.md content is delivered as a user message after the system prompt, not as part of the system prompt itself. Claude reads it and tries to follow it, but there's no guarantee of strict compliance, especially for vague or conflicting instructions.
CLAUDE.md and CLAUDE.local.md files in the directory hierarchy above the working directory are loaded in full at launch. Files in subdirectories load on demand when Claude reads files in those directories.
A symlink also works, but Windows symlinks need Administrator privileges or Developer Mode, so the @AGENTS.md import is the recommended route
Default /init reads Cursor and Copilot rules only; AGENTS.md, Devin, Windsurf and Cline rules are read only with CLAUDE_CODE_NEW_INIT=1
Subagent frontmatter carries permissionMode, hooks, mcpServers, maxTurns, memory and isolation (set to worktree to run in a temporary git worktree), and when the parent runs auto mode any permissionMode in the subagent frontmatter is ignored.
The subagent file format is markdown with YAML frontmatter, but the fields are Claude Code's own and do not map onto other tools.
4 more excerpts
Use one when a side task would flood your main conversation with search results, logs, or file contents you won't reference again: the subagent does that work in its own context and returns only the summary.
Definition precedence: managed settings, --agents CLI flag, .claude/agents/, ~/.claude/agents/, plugin agents; the definition closest to the working directory wins (v2.1.178+)
Subagent hooks support PreToolUse, PostToolUse and Stop, converted to SubagentStop at runtime
permissionMode: bypassPermissions is ignored when permissions.disableBypassPermissionsMode is set (v2.1.223+)
Claude Code docs (code.claude.com)undated (version gates reference up to v2.1.260)practitioner
'By default, if the sandbox cannot start because dependencies are missing or the platform is unsupported, Claude Code shows a warning and runs commands without sandboxing. To make this a hard failure instead, set sandbox.failIfUnavailable to true.' Native Windows is not supported.
The vendor sandbox fails open unless you configure it not to, so it cannot be the enforcement boundary.
4 more excerpts
'The operating system enforces the sandbox boundary on the running process, so it holds regardless of what the model chose to run and even if an allowed command does more than its name suggests.'
Paths and domains from both sandbox settings and permission rules are merged into the final sandbox configuration, so the sandbox is interlocked with Claude Code's own permission grammar
Protected paths (.claude settings, skills, agents, commands, hooks, .mcp.json, credentials) cannot be exempted except by disabling filesystem isolation entirely
In a linked git worktree the sandbox allows writes to the main repository's shared .git directory except hooks/ and config
Claude Code docs (code.claude.com)undated (version notes run through v2.1.248)practitioner
'Hooks are user-defined shell commands, HTTP endpoints, MCP tool calls, LLM prompts, or subagents that execute automatically at specific points in Claude Code's lifecycle.' The reference lists 33 events, five handler types and seven configuration locations, and exit 2 means a blocking error that even a JSON permissionDecision of allow cannot override.
Hooks are a vendor-specific enforcement surface with their own event vocabulary, so they belong in the per-tool adapter pile.
4 more excerpts
'For most hook events, only exit code 2 blocks the action. Claude Code treats exit code 1 as a non-blocking error and proceeds with the action, even though 1 is the conventional Unix failure code.'
Matcher semantics changed across point releases (v2.1.195, v2.1.214, v2.1.248), so hook configs are version-sensitive
The page says handlers run in the current directory with Claude Code's environment; it does not say 'full user permissions', treat that as inference
Cloud sessions on Claude Code on the web do not read local ~/.claude/settings.json; hooks there come from the repo and server-managed settings
Claude Code docs (code.claude.com)undated (version markers v2.1.172 to v2.1.239)practitioner
When using Amazon Bedrock, the /logout command is unavailable and the WebSearch tool is not available. Without pinning, model aliases such as sonnet and opus resolve to Claude Code's built-in default for Bedrock, which can lag the newest release, and Claude Code falls back to an earlier or lower-tier model at startup when the default is unavailable.
A second serving path for the same agent is a cheaper hedge than a second vendor, and it has a named feature cost.
3 more excerpts
Enabled with CLAUDE_CODE_USE_BEDROCK=1 plus AWS_REGION
A gateway that rewrites Content-Type on streaming responses forces a slower non-streaming path on every turn
'The tradeoff is that the gateway becomes infrastructure your organization operates. Claude Code adds capabilities with each release, and a gateway that doesn't forward them breaks the corresponding features.' ANTHROPIC_BASE_URL is the variable that points Claude Code at the gateway.
Routing through a gateway buys provider switching and central credentials at the price of running infrastructure that must track the agent's release cadence.
3 more excerpts
Anthropic does not endorse, maintain or audit third-party gateways and does not support routing Claude Code to non-Claude models through any gateway
Provider switching without reconfiguring machines depends on the gateway exposing a single Anthropic-format endpoint
While a gateway credential variable or apiKeyHelper is active, a developer's claude.ai subscription is not used
Claude Code docs (code.claude.com)undated (version-gated content up to v2.1.257)practitioner
Six named modes (default, acceptEdits, plan, auto, dontAsk, bypassPermissions) plus a classifier. Setting auto or bypassPermissions in .claude/settings.json or .claude/settings.local.json does not take effect; if the classifier blocks an action 3 times in a row or 20 times total, auto mode pauses, and those thresholds are not configurable.
Claude Code's permission model is a product-specific state machine that cannot be expressed in another tool's config, and parts of it cannot be set from the repo at all.
4 more excerpts
Each action goes through a fixed decision order. The first matching step wins: allow/ask/deny rules resolve immediately, read-only actions and working-directory edits auto-approve second, everything else goes to the classifier third.
On Pro, Max and Team plans the built-in starting mode is auto (v2.1.228 or later); Enterprise, API-key, claude -p, Agent SDK and Bedrock sessions start in Manual
The classifier runs on Claude Sonnet 5 by default regardless of the /model selection
Under auto mode a subagent's permissionMode frontmatter is ignored
.mcp.json uses the standard top-level mcpServers shape and accepts streamable-http as an alias for http so configurations copied from server documentation work unchanged, but an entry with a url and no type is read as a stdio server and skipped with the error 'MCP server "<name>" has a "url" but no "type"'.
The mcpServers object is the portable element; scope precedence, approval gating and the type-field gotcha are Claude Code wiring.
3 more excerpts
Scope precedence: local, project, user, plugin-provided, claude.ai connectors; same-name servers are not merged
Project-scoped .mcp.json servers need approval before first use in interactive sessions; non-interactive contexts load them without prompting unless --strict-mcp-config (v2.1.246+) or disabledMcpjsonServers is used
anthropic/requiresUserInteraction in a tool's _meta forces an approval prompt on every call even under auto or bypassPermissions (v2.1.199+)
A dev container 'is a convention rather than an enforcement boundary, because Claude Code does not require a container.' Under the built-in Bash sandbox, MCP servers and hooks are separate processes that run unconstrained on the host; the built-in Bash sandbox is the only approach Claude Code enforces itself.
Isolation that lives inside the vendor's product is a per-tool control surface; a container, VM, worktree or CI layer sits underneath every tool.
3 more excerpts
A sandboxed session that can write .mcp.json, .claude/commands or .claude/agents can persist hooks or MCP servers that run unsandboxed on the next launch
Isolation does not change what is sent to the model
Permission modes decide whether a call runs; isolation restricts what it can reach once it runs
Claude Docs (Anthropic)living docs (fetched Aug 3, 2026; inline version markers v2.1.198-v2.1.212)practitioner
'Running each Claude Code session in its own worktree means edits in one session never touch files in another.' The docs distinguish mechanisms explicitly: worktrees 'isolate file edits, while subagents and agent teams coordinate the work itself.'
File-ownership isolation is a shipped, first-class product mechanism - the enforcement layer exists today.
2 more excerpts
The page ships the enforcement mechanism but gives no guidance on how a human decides which files each session owns, and says nothing about what happens when two sessions need the same file - that absence is the gap the post addresses
Isolation is filesystem and branch level, not permission or prompt level: subagents take an 'isolation: worktree' frontmatter field, and Claude runs 'git worktree lock' while an agent is active
A change to a database system's permissions caused it to output multiple entries into a Bot Management feature file, which doubled in size and broke core traffic from 11:20 UTC, with recovery around 14:30 UTC and full normalization by 17:06 UTC.
Premise (b) of the shared-edge judgment: a shared edge provider failed the same day, in an overlapping window, from an internal change.
3 more excerpts
Names no AI vendor; the link to OpenAI's same-day incident is time-consistent but neither postmortem names the other
Explicitly not caused by an attack or malicious activity
No same-day Anthropic incident could be retrieved from the rolling status feed, so the shared-edge claim stays author's judgment
'Other exit codes - Hook failed, action proceeds (fail-open by default)'. The failClosed override ships with default false: 'When true, hook failures (crash, timeout, invalid JSON) block the action instead of allowing it through. Useful for security-critical hooks.'
A second vendor whose guardrail layer fails open by default, with the security switch shipped off.
3 more excerpts
Verification note: two independent passes disagreed on whether the page renders permission: "deny" or permission: 'deny'; that quote was dropped from the post rather than resolved
No raw-markdown endpoint (cursor.com/docs/hooks.md returns 404), so all quotes come from rendered HTML
beforeReadFile hook failures are logged and the read is allowed through
'Cursor supports AGENTS.md in the project root and subdirectories.' 'Unlike Project Rules, AGENTS.md is a plain markdown file without metadata or complex configurations.' Nested AGENTS.md files combine with parents, with more specific instructions taking precedence.
Cursor reads the shared file natively with no flip, while its .mdc rules and dashboard-managed Team Rules have no equivalent outside Cursor.
3 more excerpts
.mdc attachment logic (alwaysApply, globs, description-triggered) has no equivalent in a flat instructions file
Rules apply in the order Team Rules, Project Rules, User Rules; earlier sources win on conflict
Team and Enterprise plans can enforce rules from the Cursor dashboard so they cannot be disabled by members
Twenty-one hook events (14 supported for cloud agents plus 7 not available to them) configured in .cursor/hooks.json, and 'Exit code 2 from command hooks blocks the action (equivalent to returning permission: "deny"). This matches Claude Code behavior for compatibility.'
Cursor is the only vendor doc that states its exit-2 contract matches Claude Code's, and even so the event list is its own.
2 more excerpts
Tab hooks and workspaceOpen are excluded from cloud agents because Tab completions are an IDE feature
Config paths: <project-root>/.cursor/hooks.json and ~/.cursor/hooks.json
Permissions are permissions.allow and permissions.deny arrays of rule strings such as Shell(rm), Read(.env*), Write(**/*.key) and Mcp(server:tool) in .cursor/cli.json or ~/.cursor/cli-config.json, and 'Deny rules take precedence over allow rules.'
Cursor's permission grammar is its own syntax in its own files, with no cross-tool equivalent.
2 more excerpts
Shell(commandBase) matches on the first token of the command line
Without an allowlist entry, each WebFetch prompts for approval
Cursor reads subagent files from .claude/agents/ (Claude compatibility) and .codex/agents/ (Codex compatibility) in addition to .cursor/agents/, with .cursor/ taking precedence on name conflicts; its own frontmatter fields readonly and is_background have no equivalent elsewhere.
Subagent non-portability is field-level only: the file location is interoperable, the fields (permissionMode, hooks, mcpServers on Claude; readonly, is_background on Cursor; TOML developer_instructions on Codex) are not.
2 more excerpts
Correction: the earlier claim that Cursor subagents do not map onto Claude Code or Codex was broken by this page; Cursor reads .claude/agents/ and .codex/agents/ explicitly
Project MCP config lives in .cursor/mcp.json and global config in ~/.cursor/mcp.json, using the mcpServers shape, with Cursor-specific interpolation tokens such as ${env:NAME}, ${userHome}, ${workspaceFolder} and ${pathSeparator}.
The mcpServers object is the portable element; the file locations, interpolation and enterprise controls are Cursor wiring.
2 more excerpts
Enterprise Team Settings add command and URL entries, tool allowlists and network modes (Allow all, Allowlist, Deny all, No sandbox)
Erik Schluntz, Barry Zhang (Anthropic)December 19, 2024practitioner
Five named workflow patterns (Prompt chaining, Routing, Parallelization, Orchestrator-workers, Evaluator-optimizer), each with its own section and diagram. The only intermediate control point diagrammed as a node is a programmatic 'gate'; human checkpoints appear in a single line of prose.
The topology primitives predate later attempts to rename them, and no frontier-lab source elevates the human checkpoint to a named, diagrammed pattern the way it does the programmatic gate.
2 more excerpts
The subtractive bar, with 'only' italicized on the page: 'you should consider adding complexity only when it demonstrably improves outcomes.'
Anthropic gates 'multi-step agentic systems' while still endorsing multi-step workflows for well-defined tasks.
The context.fileName setting in settings.json ships an example list of 'AGENTS.md, CONTEXT.md, GEMINI.md', and large files can import others with the @file.md syntax using relative or absolute paths.
Gemini CLI is a one-key flip to read the shared AGENTS.md.
3 more excerpts
The page banner reads: 'Unpaid tier and Google One users: Gemini CLI was replaced by Antigravity CLI on June 18th, 2026', the same date as the last-updated stamp
Three-tier load order: global ~/.gemini/GEMINI.md, workspace ancestors, and just-in-time loading when a tool touches a directory
Gemini CLI itself is not durable for unpaid-tier readers; the post has to address that
Eleven hook events across four groups: BeforeTool and AfterTool; BeforeAgent and AfterAgent; BeforeModel, BeforeToolSelection and AfterModel; SessionStart, SessionEnd, Notification and PreCompress. None shares a name with a Claude Code event.
Gemini CLI's hook vocabulary is disjoint from Claude Code's, so hooks are rewritten rather than copied.
3 more excerpts
The count of 11 is a manual tally across four section headings; there is no single master table
Blocking semantics differ per hook type: BeforeAgent aborts the turn and erases the prompt, AfterAgent rejects the response and triggers a retry
Configured in layered settings.json files (project .gemini/, user ~/.gemini/, system /etc/gemini-cli/)
A priority-ordered rule engine: 'When a large language model wants to execute a tool, the policy engine evaluates all rules to find the highest-priority rule that matches the tool call', with allow, deny and ask_user decisions loaded from user, workspace (currently disabled) and admin TOML tiers, under an approval hierarchy of plan < default < autoEdit < yolo.
Gemini CLI's permission model is structurally unlike Codex's two flags and Claude Code's six modes, so there is no field-level translation.
3 more excerpts
ask_user is treated as deny in non-interactive mode
Workspace policies at $WORKSPACE_ROOT/.gemini/policies/*.toml are currently disabled; admin policies live in OS-specific system paths
Allow-for-all-future-sessions includes the current mode and all more permissive modes in the hierarchy
Custom agents are Markdown files with YAML frontmatter in .gemini/agents/ or ~/.gemini/agents/, with name and description required and fields including kind, tools, mcpServers, model, temperature, max_turns (default 30) and timeout_mins (default 10); the markdown body becomes the system prompt.
Same file shape as Claude Code's subagents, different directory and schema, so the definition is re-authored per tool.
3 more excerpts
model: inherit is not a documented value on this page
Subagents are treated as virtual tool names for policy matching, so access is governed by the TOML policy engine rather than by frontmatter
'Gemini CLI uses the mcpServers configuration in your settings.json file to locate and connect to MCP servers.' The trust option 'bypasses all tool call confirmations for this server (default: false)' and the docs say to use it cautiously and only for servers you completely control.
The mcpServers block is the same shape across tools, while trust and allowlist merge rules are Gemini-specific.
3 more excerpts
The literal .gemini/settings.json path is not returned verbatim from this page; agents.md states it
excludeTools takes precedence over includeTools; when two sources both give an allowlist, only tools in both lists are enabled
Transports: stdio, SSE and streamable HTTP; fields include cwd, httpUrl, headers, timeout, targetAudience and targetServiceAccount
The vendor's own manual equivalent of its experimental --worktree flag is 'git worktree add ../project-feature-search -b feature-search' followed by cd and gemini; the flag stores worktrees under .gemini/worktrees/ and does not automatically delete the worktree or branch.
Worktree isolation is plain git, so it belongs in the layer that survives a vendor swap.
3 more excerpts
Gated behind an experimental opt-in: "experimental": { "worktrees": true } in settings.json
Cleanup is manual: git worktree remove .gemini/worktrees/<name> --force
The page carries the Gemini CLI to Antigravity CLI replacement banner
GitHub (anthropics/claude-code issue 6235)Opened Aug 21, 2025; closed Aug 17, 2026practitioner
The year-long request (6,592 reactions, 389+ comments) was closed as completed by a comment pointing at the workaround: 'create a CLAUDE.md containing just @AGENTS.md (an import), or symlink CLAUDE.md to AGENTS.md.' A thread report from 2026-07-04 says a recent update blocked symlinked file updates with 'cannot write through symlinks'.
Native AGENTS.md support in Claude Code was declined in favor of the import, and the symlink route has broken in the field, which is why the post recommends import over symlink.
3 more excerpts
Closed with state_reason completed, not 'not planned' and not an inactivity auto-close
Maintainer status of the closing account (bcherny) was not independently verified beyond authoring the closing response
Community pushback the same night: the linked doc still requires a CLAUDE.md, a workaround rather than adoption
CLAUDE.md files in subdirectories are not being automatically loaded when accessing files in those directories, contrary to what the documentation states.
The lazy tier has unresolved reliability reports — documented design, verify on your surface (single macOS report, closed unresolved).
GitHub (anthropics/claude-code)opened Feb 11, 2026; closed not-planned Mar 22, 2026practitioner
only the root-level CLAUDE.md is loaded at session start, and no subdirectory CLAUDE.md files are ever injected — even after multiple Read tool calls into those directories.
Lazy loading failed on the VS Code extension across three versions (2.1.39/2.1.45/2.1.49; CLI reportedly fine) — surface-specific reliability caveat.
GitHub (anthropics/claude-code)opened Nov 16, 2025; closed not-plannedpractitionerpartial
A modular CLAUDE.md structure with 6 referenced files (~2,100 lines total) consumes the same tokens as a monolithic file, providing organizational benefits only.
A user measured the @imports split delivering zero token savings, and Anthropic declined the lazy-imports request (author-self-reported measurement, not maintainer-confirmed).
1 more excerpt
85-90% of loaded content is irrelevant to most conversations
'If multiple rulesets target the same branch or tag in a repository, the rules in each of these rulesets are aggregated' and 'the most restrictive version of the rule applies.' 'Anyone with read access to a repository can view its active rulesets.'
Enforcement on the host survives any agent swap because no vendor config can reach it, and it is auditable without admin access.
3 more excerpts
'Evaluate mode' does not appear on this page; do not cite it to this URL
Up to 75 rulesets per repository and 75 organization-wide
Bypass can be granted to roles, teams or GitHub Apps when a ruleset is created; rulesets and branch protection rules enforce alongside each other
'By default, the restrictions of a branch protection rule don't apply to people with admin permissions to the repository or custom roles with the bypass branch protections permission' until 'Do not allow bypassing the above settings' is enabled; a required status check pinned to an app blocks merging 'if the status is set by any other person or integration.'
Two settings decide whether the host gate actually holds: closing the admin bypass and pinning the required check to its app.
3 more excerpts
Covers legacy branch protection rules; rulesets have their own bypass-actor model
Only one branch protection rule applies at a time, a restriction that does not apply to rulesets
Required checks must have a successful, skipped or neutral status before changes land
Google (Gemini CLI docs)undated, main branch (fetched Jul 27, 2026)practitioner
'For global rules (those without an argsPattern), tools that are denied are completely excluded from the model's memory.' The model never sees the tool as an option.
A deny decision is not a request to the model - it is a change to what exists. Deterministic policy evaluation as reviewable code.
4 more excerpts
SCOPE: the memory-exclusion claim applies only to global rules without an argsPattern; the doc says nothing about argument-conditional deny rules
Precedence is arithmetic: final_priority = tier_base + (toml_priority / 1000), tier bases Default 1 through Admin 5 - tier always dominates because the fractional term can never reach 1
'The first rule that matches determines the outcome' - first in priority order, not file order
In non-interactive mode ask_user is treated as deny
Google Developers BlogMay 19, 2026 (transition effective Jun 18, 2026)practitionerpartial
'On June 18, 2026, Gemini CLI and Gemini Code Assist IDE extensions will stop serving requests for Google AI Pro and Ultra, as well as those using it free of charge using Gemini Code Assist for individuals.'
The one real vendor swap in the record: a vendor retired its own tool out from under individual users.
3 more excerpts
The blog names Google AI Pro and Ultra plus free individuals, not 'Google One'; that phrasing is the geminicli.com banner's
Organizations on Gemini Code Assist Standard or Enterprise licenses, or using Gemini Code Assist for GitHub through Google Cloud, keep unchanged access
Gemini Code Assist for GitHub also stopped new installations on June 18, 2026
yes, compaction and smaller models help on cost per step. But my issue wasn't just inefficiency, it was agents retrying when they shouldn't. I needed visibility + limits per agent/task, and the ability to cut it off, not just optimize it.
Practitioners want per-agent/per-task limits and a hard cut-off, not just cost optimization — the wedge is attribute-and-enforce, not optimize.
4 more excerpts
My AGENTS.md is 845 lines and it only started getting good once it got that long" (Sammi), directly contested by "sweet spot is between 60 and 120 lines. With psuedo xml tags between sections" (typpilol)
Budget alerts are not a kill switch. Credits are not protection.
Claude often ignores CLAUDE.md / The more information you have in the file the more it gets ignored
"cost control is a policy problem - we certainly don't need to use opus 4.6 for a simple test refactor... we need a way to measure cost / performance for agents on individual repos, with individual types of tasks..." (author bisonbear, id 47563774)
Harrison Chase (LangChain)June 16, 2025practitioner
'The key insight is that read actions are inherently more parallelizable than write actions.' Parallelizing writes creates a dual problem: communicating context between agents and then merging their outputs coherently.
Whether an edge is safe to draw depends on whether the work crossing it reads or writes, which is why research-shaped multi-agent systems succeed where coding-shaped ones struggle.
2 more excerpts
Explicitly reconciles Cognition's 'Don't Build Multi-Agents' with Anthropic's research-system post as two correct answers for different work: 'Despite their opposing titles, I would argue they actually have a lot in common.'
Reproduces Anthropic's four-field subagent spec as an attributed block quote, so it is commentary on that material rather than independent corroboration of it.
Ivan Kahl / DometrainJanuary 15, 2026practitionerpartial
You cannot craft the perfect CLAUDE.md file immediately. Instead, treat it as a living document.
CLAUDE.md is a living document refined over time, not a one-shot artifact.
2 more excerpts
Claude Code agents have a context window, and the CLAUDE.md file gets added to the agent's context. Any unnecessary instructions and wordy sentences will consume more of that context.
Always review the CLAUDE.md file and correct any assumptions or missing details related to project architecture.
In 56% of eval cases the skill was never invoked. A compressed 8KB docs index embedded directly in AGENTS.md achieved a 100% pass rate while skills maxed out at 79%; skills without explicit instructions matched the no-docs baseline at 53%.
Whether an agent uses a portable artifact depends on the harness deciding to load it, which is a per-tool behavior rather than a property of the file.
2 more excerpts
Vendor eval on Next.js 16 API tasks; no inference-cost delta is reported on the page
The docs injection started at about 40KB and was compressed to 8KB, an 80% reduction
Frontier thinking LLMs can follow ~ 150-200 instructions with reasonable consistency.
There is a practical instruction ceiling — even frontier models only follow roughly 150-200 instructions consistently — so every line in CLAUDE.md competes for a finite budget.
4 more excerpts
At HumanLayer, our root CLAUDE.md file is less than sixty lines.
Claude Code's system prompt contains ~50 individual instructions
Smaller models get MUCH worse, MUCH more quickly
LLMs bias towards instructions that are on the peripheries of the prompt
Lance Martin, Gabe Cemaj, Michael Cohen (Anthropic Engineering)Apr 8, 2026practitioner
'Our only window in was the WebSocket event stream, but that couldn't tell us where failures arose, which meant that a bug in the harness, a packet drop in the event stream, or a container going offline all presented the same.'
A first-party operator account that having a telemetry stream is not the same as being able to localize a failure.
4 more excerpts
getEvents(), allows the brain to interrogate context by selecting positional slices of the event stream
Cited for the BEFORE state only. A full-text search confirms the article never reports an observability after-state
It names decoupling as the fix and scopes its durable event log to crash recovery and model-side context interrogation
Its quantified improvements (p50 TTFT down roughly 60%, p95 over 90%) are latency, not observability
The pause is implemented on the graph's own persistence layer: 'Every step of the graph, it reads from and then writes to a checkpoint of that graph state', which is what allows the run to 'pause execution of the graph half way through, and then resume after some time'.
The human checkpoint is durable because it rides the graph's existing checkpoint mechanism rather than an external callback, which is what makes it structurally different from a manual review step in a runbook.
1 more excerpt
Named human-in-the-loop patterns: Approve or Reject, Review and Edit State, Review Tool Calls, and multi-turn conversation in a multi-agent setup.
'The interrupt function pauses graph execution and returns a value to the caller. When you call interrupt within a node, LangGraph saves the current graph state and waits for you to resume execution with input.' Named patterns include approval workflows, review and edit state, interrupts in tools, and validating human input.
A human checkpoint can be a first-class graph node with durable state rather than a process stage described in prose - the mechanism ships today in a production framework.
3 more excerpts
Durability is explicit: state is saved by the checkpointer, which in production should be database-backed.
Idempotency gotcha: 'The node restarts from the beginning... so any code before the interrupt runs again.'
The documentation contains no notion of review capacity, cost, or how many checkpoints is too many.
'The evaluator and permission control should likely sit outside the loop that evolves harness, with held-out tests, trace audits, and human review at decision points that matter.'
Enforcement belongs outside the loop it governs - and, because harnesses evolve, the audit is recurring rather than one-time.
4 more excerpts
SCOPE: scoped specifically to self-modifying harnesses, not agent harnesses generally; the 'should likely' hedge is the author's own
'A harness is the system surrounding a base model that orchestrates execution and decides how the model thinks and plans, calls tools and acts, perceives and manages context, stores artifacts, and evaluates results'
The words deterministic and probabilistic appear nowhere in the post - do not attribute that framing to her
'The core interface of mainstream coding agents has become stabilized across Claude Code, Codex, OpenCode, and Cursor-style agents'
Molisha Shah (Augment Code)March 16, 2026 (updated June 18, 2026)practitionerpartial
Six patterns for parallel coding agents, including git-worktree isolation (each agent gets its own working directory and index while sharing one .git object database) and sequential merges. 'Git detects textual conflicts, not semantic ones. Authoritative guidance still converges on mandatory human review for logic-level contradictions.'
Mechanical isolation between parallel coding agents solves textual collision but not semantic conflict, which is what leaves a human judgment step irreducible at the merge boundary.
2 more excerpts
Vendor guide authored by a go-to-market staff member rather than an engineer; the technical core is real but a '40% reduction in hallucinations' claim about the vendor's own product is unverified marketing and is not cited.
Humans are given a tier in a review pipeline ('Reserve humans for semantic correctness and architecture'), not a position in a graph.
OpenAIundated, on or before Feb 11, 2026 (Wayback-bounded)practitionerpartial
'Agents are most effective in environments with strict boundaries and predictable structure, so we built the application around a rigid architectural model.'
The canonical build-a-harness text - and the differentiation foil: written from an empty repository, with zero coverage of inherited vendor defaults.
4 more excerpts
Direct fetch returns HTTP 403 and web.archive.org is blocked to the tool; content read via a text-extraction proxy
Page carries no byline and no publication date - attribute to OpenAI, not to an individual
Its 'boundaries' are architectural layer-dependency lint rules inside the application codebase, not harness control surfaces
Posture cuts against gates: 'The repository operates with minimal blocking merge gates'
OpenAI's stated default: 'Start with one agent whenever you can. Add specialists only when they materially improve capability isolation, policy isolation, prompt clarity, or trace legibility. Splitting too early creates more prompts, more traces, and more approval surfaces without necessarily making the workflow better.'
Both major labs publish a single-agent default with an explicit justification bar, and OpenAI counts human approval surface as one of the costs that premature decomposition imposes.
3 more excerpts
'Approval surfaces' is used once and never defined, on this page or the parent guide.
Neither page uses 'review capacity' or 'reviewer bandwidth', and neither frames human review as a binding constraint on agent count - that framing belongs to the post, not to OpenAI.
'Trace legibility' is listed as a reason to split, which cuts against reading this as a blanket anti-splitting stance.
'Some specialized tool paths can opt out of the default hook path. Treat tool hooks as a useful guardrail, not a complete enforcement boundary.'
A vendor stating in its own documentation that the layer most teams build policy on is not an enforcement boundary.
4 more excerpts
Codex hooks fire at 12 lifecycle points (PreToolUse, PermissionRequest, PostToolUse, PreCompact, PostCompact, UserPromptSubmit, SubagentStop, Stop, Interrupt, SessionStart, SubagentStart, SessionEnd), are configured in .codex/hooks.json or .codex/config.toml, and 'you can also use exit code 2 and write the blocking reason to stderr.'
The named hole: 'Hosted tools, such as WebSearch... don't use the local function-tool hook path'
Codex also fails open: a PreToolUse hook returning unsupported fields is marked failed and 'continues the tool call'
'Multiple matching command hooks for the same event are launched concurrently, so one hook can't prevent another matching hook from starting'
OpenAI (status.openai.com)Jun 9 to Jun 10, 2025practitionerpartial
ChatGPT users experienced elevated error rates reaching about 35% at peak while API error rates peaked near 25%, and API availability dropped to 75%; the absence of break-glass tooling to rapidly restore network connectivity on affected nodes extended the overall recovery timeline.
The second vendor has its own outage character: a routine host-OS update on GPU nodes fanned out into a roughly 15-hour cross-product outage.
3 more excerpts
The write-up does not mention Codex anywhere; only ChatGPT and the API are named, so the earlier 'Codex listed as affected' claim was dropped
Root cause was a systemd-networkd restart conflicting with a production networking agent, which removed all routes from affected nodes
Remediation disabled automatic daily updates on GPU systems
'We have confirmed that the incident is caused by an issue with one of our third-party service providers.' The outage affected APIs and ChatGPT from 12:12 PM to 3:18 PM, roughly 3 hours 6 minutes.
Premise (a) of the shared-edge judgment: OpenAI attributed a multi-hour outage to an unnamed upstream provider.
3 more excerpts
No provider is named; do not assert Cloudflare from this source
The promised root cause analysis was not located
Over an hour passed before the cause category was named
The import maps instruction files to AGENTS.md, settings.json to config.toml, slash commands to Skills, subagents to Codex subagents, MCP configuration to Codex MCP configuration and hooks to Codex hooks, then lists tool restrictions and permissions, MCP auth and headers, hook behavior, plugins and argument-bearing prompts as items to review manually.
The vendor's own migration tool draws the line between what carries over and what has to be rebuilt per tool.
3 more excerpts
Supported sources: Claude Code, Claude Cowork and Cursor for the desktop app; Claude Code and Cursor for Codex CLI
The page describes importing into Codex only, never exporting; use 'items to review', not 'silently converted'
Chats import at most 50 from the last 30 days, which is state rather than configuration
Codex checks each directory in this order: AGENTS.override.md, AGENTS.md, TEAM_GUIDE.md, .agents.md (fallbacks set via project_doc_fallback_filenames), uses only the first non-empty file per level, and stops adding files once the combined size reaches project_doc_max_bytes, default 32 KiB.
Codex reads AGENTS.md natively and rebuilds the instruction chain root-down each run, so no flip is required.
2 more excerpts
Files closer to the current directory override earlier guidance because they appear later in the chain
Spot-check the literal fallback example array before quoting
Permissions are a cross-product of approval_policy (untrusted, on-request, never, or granular) and sandbox_mode (read-only, workspace-write, danger-full-access), where danger-full-access disables sandboxing entirely; the combined --dangerously-bypass-approvals-and-sandbox flag is aliased --yolo.
Codex's two-axis permission model has no field-level equivalent to Claude Code's six named modes, so permissions do not port.
2 more excerpts
Codex loads project-scoped .codex/ config layers only when the project is trusted and may start read-only until the working directory is trusted
The six named sandbox-and-approval combinations come from the companion agent-approvals-security page
Every standalone custom agent file must define name, description and developer_instructions, lives as TOML under .codex/agents/ or ~/.codex/agents/, and the docs say 'the format may evolve as authoring and sharing mature.'
The Codex subagent format is TOML with its own required keys, not the markdown-plus-frontmatter shape Claude Code, Gemini CLI and Cursor share.
3 more excerpts
Codex reapplies the parent turn's live runtime overrides (such as /permissions changes or --yolo) when spawning a child, even if the agent file sets different defaults
Settings resolve from an explicit spawn value, then the [agents] default, then the parent's value
Other config.toml keys such as model, sandbox_mode, mcp_servers and skills.config can be included in an agent file
MCP servers are configured as [mcp_servers.<server-name>] TOML tables in ~/.codex/config.toml or a project-scoped .codex/config.toml (trusted projects only), and 'the ChatGPT desktop app, Codex CLI, and IDE extension share this configuration.'
Codex shares MCP config across OpenAI's own surfaces only, in TOML rather than the JSON mcpServers shape the other three tools use.
3 more excerpts
Per-server default_tools_approval_mode accepts auto, prompt, writes and approve, plus per-tool approval_mode
Codex validates any returned iss before exchanging an OAuth authorization code; a mismatch always rejects
Codex reads the MCP instructions field returned at initialization as server-wide guidance
Codex creates worktrees in $CODEX_HOME/worktrees in a detached HEAD state, keeps the most recent 15 Codex-managed worktrees by default, and copies ignored files only if they match .worktreeinclude (AGENTS.override.md is copied automatically).
Vendor-managed worktrees are a convenience over plain git worktree that is scoped to one product's parallel-chat feature.
2 more excerpts
The mechanism is scoped to the Codex desktop app running parallel chats
AGENTS.override.md is a Codex-specific override file distinct from the AGENTS.md convention
'Pair AGENTS.md with infrastructure that enforces those rules: pre-commit hooks, linters, and type checkers catch issues before you see them.' Files closer to the working directory take precedence, and nested AGENTS.md supersede parents while the repo overrides ~/.codex/AGENTS.md.
The vendor's own advice is that the instruction file is not the enforcement layer; enforcement lives in tooling no agent can talk past.
2 more excerpts
Skills live at ~/.agents/skills (global) and .agents/skills (repo) as SKILL.md with optional scripts, references and assets
agents/openai.yaml (MCP dependencies for skills) and Plugins are explicitly OpenAI/Codex-specific
Simon Willison16th October 2025practitionerpartial
This is _very_ token efficient: each skill only takes up a few dozen extra tokens, with the full details only loaded in should the user request a task that the skill can help solve.
Named-author validation of metadata-first, on-demand loading (about Skills; the CLAUDE.md parallel is the post's).
'A coding agent is a piece of software that acts as a harness for an LLM, extending that LLM with additional capabilities that are powered by invisible prompts and implemented as callable tools.'
The harness's prompts are invisible to the user by construction - the closest practitioner framing to the unreviewed-defaults claim.
1 more excerpt
'A tool is a function that the agent harness makes available to the LLM'
'you can try telling it not to in your own prompt, but how confident can you be that your protection will work every time?'
A system-prompt instruction is not a security control - the cleanest one-line statement of the post's premise.
1 more excerpt
On vendor guardrails claiming '95% of attacks': 'in web application security 95% is very much a failing grade' - note the 95% is Willison characterizing vendor marketing, not his own measurement
context engineering is the delicate art and science of filling the context window with just the right information for the next step...task descriptions and explanations, few shot examples, RAG, related (possibly multimodal) data, tools, state and history
Context engineering, not prompt engineering, is the real discipline: filling the window with the right information environment for the next step.
1 more excerpt
the art of providing all the context for the task to be plausibly solvable by the LLM
Towards Data Science (Mostafa Ibrahim)March 20, 2026practitioner
The agent optimises locally. At each step, it asks, 'Do I have enough?' and when the answer is uncertain, it defaults to 'get more'. Without hard stopping rules, the default spirals.
Without a hard stop rule, an agent's local 'get more' default turns retrieval into an unbounded budget fire; capping cycles and abstaining is the control.
3 more excerpts
Three cap retrieval cycles. After three failed passes, return a best-effort answer with a confidence disclaimer.
agents making 200 LLM calls in 10 minutes, burning $50–$200 before anyone noticed
costs spike 1,700% during a provider outage as retry logic spiralled out of control
The honest column in the ledger: JIT buys these at the price of retrieval latency on the steps that load (usually trivial next to a model call, but nonzero), a new failure mode (an unresolvable reference must surface as an honest error, not a hallucinated payload), and a dependency on description quality — the agent loads from the catalog's one-liners, so a bad stub hides a good payload.
JIT context introduces two specific liabilities leaders must design for: unresolvable references must fail loud as honest errors, and retrieval quality is capped by the quality of catalog descriptions.
3 more excerpts
in a loop, the window is re-sent every step, so a preloaded handbook isn't one payment but thirty
a preloaded copy is a snapshot that ages as the run proceeds, while a reference resolves to the current state of the file, the ticket, the database at the moment of use
long-context research and practitioner experience agree that models degrade as windows fill with low-relevance text
This is powerful for keeping your main file lean. Put detailed instructions in separate markdown files, then reference them. Claude pulls in the content when relevant.
The SERP's reputable competitor teaches the naive @imports model the docs refute — the foil for the kicker.
'A harness is every piece of code, configuration, and execution logic that isn't the model itself. A raw model is not an agent. But it becomes one when a harness gives it things like state, tool execution, feedback loops, and enforceable constraints.'
The canonical definition, from the coinage. Cite Trivedy rather than downstream restatements.
1 more excerpt
The component enumeration explicitly places System Prompts inside the harness alongside tools, sandbox, orchestration logic, and hooks - so 'harness vs CLAUDE.md' is not the boundary the field draws
negative instructions can be unreliable as user prompts
Negative 'don't do that' rules are unreliable in a user message like CLAUDE.md, so positive, runnable framing is preferable — reserving DO-NOT for hard safety boundaries.
3 more excerpts
Reddit user reported Claude Code created duplicate files despite explicit 'NEVER create duplicate files' rule
Gemini models have 'hit-or-miss' performance with negative commands
They are effective at preventing unethical or harmful behavior, especially when used in system prompts
Knowing an agent’s output is actually correct, beyond a green build.
Albayaydh, Zhao, FlechaisJul 7, 2026data
A synthesis of 27 benchmark, taxonomy, and audit papers spanning 19 benchmarks finds that additional scaffolding does not consistently improve reliability.
The honesty brake on 'more harness is better' - and the reason this post's claim is bounded to blast radius rather than quality.
3 more excerpts
Failures compound nonlinearly with task length
Strong performance on individual sub-tasks does not reliably translate into end-to-end success
Secondary synthesis - every number in it is someone else's measurement
32.67% of successful SWE-bench patches involved solution leakage (the fix present in the issue report or comments) and 31.08% passed on weak tests; filtering both drops SWE-Agent+GPT-4's resolution rate from 12.47% to 3.97%.
A third of measured SWE-bench success was answer leakage, a concrete mechanism by which leaderboard scores inflate without capability.
1 more excerpt
Over 94% of benchmark issues predate LLM knowledge cutoff dates
Andrew Gelman, Eric Loken (Columbia University / Penn State)November 14, 2013data
Researcher degrees of freedom can produce a multiple-comparisons problem even where researchers perform only a single analysis on their data, because the analysis actually run is one of many that would have been equally reasonable had the data come out differently.
A team reading its own agent-PR dashboard once, in good faith, with a hypothesis stated in advance, is already exposed to the fragility — no p-hacking or repeated analysis is required for the result to be contingent on arbitrary choices.
2 more excerpts
'The researcher degrees of freedom do not feel like degrees of freedom because, conditional on the data, each choice appears to be deterministic'
Cited from the 2013 unpublished manuscript; the 2014 American Scientist version ('The Statistical Crisis in Science') was not reachable for verification and is not the version quoted
Teams delay building evals thinking they need hundreds of tasks; in reality 20-50 simple tasks drawn from real failures is a great start, structured by task/trial/outcome vocabulary.
A working internal agent eval suite is a 20-50 task project, not an infrastructure program - removing the main excuse for deciding from public leaderboards instead.
4 more excerpts
Opus 4.5 initially scored 42% on CORE-Bench; after fixing grading bugs and using a less constrained scaffold, the same model's score jumped to 95%.
So as not to unnecessarily punish creativity, it's often better to grade what the agent produced, not the path it took.
'With frontier models, a 0% pass rate across many trials (i.e 0% pass@100) is most often a signal of a broken task, not an incapable agent'
The initial 42% observation is Anthropic citing an external report; the diagnosis and the 95% re-run are Anthropic's own
In internal experiments spanning six compute-resource configurations on GKE with model, harness, and task set held fixed, the gap between the most- and least-resourced setups on Terminal-Bench 2.0 was 6 percentage points (p < 0.01); infrastructure error rates fell from 5.8% under strict 1x enforcement to 2.1% at 3x headroom and 0.5% uncapped.
Infrastructure configuration alone produces score differences exceeding the few-point margins that separate top leaderboard entries, so cross-infrastructure leaderboard comparisons are not decision-grade evidence for a model swap.
4 more excerpts
Infrastructure configuration can swing agentic coding benchmarks by several percentage points - sometimes more than the leaderboard gap between top models. The gap between most- and least-resourced setups on Terminal-Bench 2.0 was 6 percentage points (p < 0.01).
A 2-point lead on a leaderboard might reflect a genuine capability difference, or it might reflect that one eval ran on beefier hardware, or even at a luckier time of day, or both.
Top leaderboard spots are often separated by just a few percentage points, per the post's own framing
Resource headroom is an eval-infrastructure design requirement: error rate falls an order of magnitude from strict to uncapped provisioning
Their hyper-productivity is revealing a significant 'speed vs. trust' gap. Recent, deeper examinations of agent-generated code and agent-driven PRs reveal that a large percentage of agent efforts fail to meet the quality bar of being truly 'merge-ready,' often containing subtle regressions, superficial fixes, or a general lack of engineering hygiene.
Agent hyper-productivity creates a speed-vs-trust gap where most agent PRs aren't merge-ready, overwhelming review capacity.
3 more excerpts
29.6% of 'plausible' fixes introduced behavioral regressions or were incorrect upon rigorous retesting
True solve rates for GPT-4 patches dropped from 12.47% to 3.97% after detailed manual audits
Over 68% of agent-generated pull requests reportedly face long delays or remain unreviewed, creating an urgent need for scalable review automation.
arXiv (Sabrina Haque, Sarvesh Ingale, Christoph Csallner)Submitted January 7-8, 2026datapartial
Across agents, test-containing PRs are more common over time and tend to be larger and take longer to complete, while merge rates remain largely similar.
Whether an agent PR includes tests varies and doesn't correlate with merge outcomes, so test presence is a signal to read, not proof of quality.
2 more excerpts
We observe variation across agents in both test adoption and the balance between test and production code within test PRs
Testing is a critical practice for ensuring software correctness and long-term maintainability
arXiv (Yue Liu, Ratnadira Widyasari, Yanjie Zhao, Ivana Clairine Irsan, Junkai Chen, David Lo)Submitted 30 March 2026 (v2 revised 26 April 2026)datapartial
22.7% of tracked AI-introduced issues still survive at the latest version of the repository. These findings show that AI-generated code can introduce long-term maintenance costs into real software projects.
Over a fifth of AI-introduced issues survive at HEAD, so AI code accrues durable technical debt at scale unless verification catches it.
3 more excerpts
302.6k verified AI-authored commits from 6,299 GitHub repositories
more than 15% of commits from every AI coding assistant introduce at least one issue
code smells are by far the most common type" / "89.3% of all issues
Berkeley RDI (Hao Wang, Qiuyang Mang, Alvin Cheung, Koushik Sen, Dawn Song)April 2026data
We built an automated scanning agent that systematically audited eight among the most prominent AI agent benchmarks — SWE-bench, WebArena, OSWorld, GAIA, Terminal-Bench, FieldWorkArena, and CAR-bench — and discovered that every single one can be exploited to achieve near-perfect scores without solving a single task.
Every major agent benchmark can be gamed to near-perfect scores without solving anything, so self-reported benchmark performance is structurally untrustworthy.
4 more excerpts
Every one of eight major AI agent benchmarks audited can be exploited to near-perfect scores without solving a single task - including 100% on SWE-bench Verified via a 10-line conftest.py that hooks pytest and rewrites every test result to passed.
A conftest.py file with 10 lines of Python 'resolves' every instance on SWE-bench Verified.
SWE-bench Verified (500 tasks) — 100% score via pytest hooks
Benchmark scores are actively being gamed, inflated, or rendered meaningless, not in theory, but in practice.
Applying rank confidence intervals to MMLU abstract-algebra rankings, the authors conclude the observed ranking cannot be trusted and all models are statistically interchangeable.
Once uncertainty is displayed, observed leaderboard orderings among top models frequently collapse into statistical ties.
On MMLU (57 subjects), three distinct models can be ranked as the fourth from the top - statistically consistent with the same rank position - with uncertainty driven more by between-subject variability than prompt variants.
Leaderboard ranks published as single values hide that several models are often statistically tied for the same position.
Carlo A. Furia, Richard TorkarJanuary 28, 2025 (latest revision April 1, 2026)datapartial
In a study of 105 projects (45 Java, 60 Python) and 952 developers, omitting programmer skill inflated the measured programming-language effect on code quality roughly fourfold, from an adjusted -0.012 to a confounded -0.052. Sensitivity analysis showed a scaled-mean difference of -0.062 would be sufficient to flip the sign of the measured effect.
An aggregate software-engineering trend can be substantially distorted — and, at plausible confounding strengths, sign-flipped — by omitting a variable that is a real determinant of the outcome, which is what a dashboard that does not condition on team properties is doing.
3 more excerpts
'Additional data about X and Y (i.e., sampling more datapoints) is not going to help; in fact, it may just entrench our reliance on the biased estimate by reducing its variance and giving the false impression of reliability'
The sign flip is a sensitivity-analysis threshold showing how little confounding would suffice, not an observed reversal in the data; neither case study concerns code review
Preprint with no journal-ref; Simpson's paradox is not discussed anywhere in the paper
Christopher Kelly, Angelica Chowdhury, Alexandra Campili, Bimpe Ayoola, Devin Barbour, Thomas Chen Dawson, Ze Shen Chin, Rokas GipiškisMay 3, 2026 (revised July 26, 2026; ICML 2026 Workshop on Technical AI Governance)data
Synthesizes five principles from established validity frameworks into 33 guidelines for AI evaluation trials, naming construct underrepresentation and construct-irrelevant variance as the failure modes that make AI-effect measurements uninterpretable.
Where construct validity is weak, a study can show that scores increased with AI without credibly claiming the underlying capability improved — which is why 'is AI good for code review' is not merely contested but malformed as a measurable question.
2 more excerpts
Guideline 18 is directly applicable to reading a vendor claim: 'Apply identical quality rubrics, performance benchmarks, and scoring criteria to human-only and human+AI outputs'
Software engineering appears as a donor methodology field and 'coding competence' as an example construct; the paper analyzes no code-review study
Clustered standard errors on public evals can be over 3X larger than naive standard errors, and detecting an absolute score difference of 0.03 at 80% power requires an eval of at least ~969 independent questions.
Most reported model-to-model benchmark gaps are narrower than honestly computed confidence intervals - the statistical foundation for treating small leaderboard gaps as ties.
2 more excerpts
The same pair of models can differ significantly on one benchmark (MATH) and not on others (HumanEval, MGSM) in the paper's worked example
The Llama 3 paper's reported confidence intervals are judged likely anti-conservative (too narrow)
A good scaffold can increase SWE-bench Verified performance by up to 20%, so scores reflect the sophistication of the scaffold as much as the capability of the underlying model.
Scaffold quality is a confound baked into every SWE-bench Verified score - the leaderboard measures a model-plus-scaffold system.
Across five runs, run-to-run standard deviations were 2.0 percentage points for SWE-Doctor (the most stable agent), 2.2 for mini-SWE-agent, and 3.4 for live-SWE-agent; SWE-Doctor's Pass@5 was 70.0% against All@5 of 40.0%.
Even the most stable SWE-bench-family agents swing multiple percentage points between identical runs - variance comparable to the gaps separating leaderboard leaders.
1 more excerpt
The 30-point spread between Pass@5 (solves at least once) and All@5 (solves every time) is its own nondeterminism exhibit
Under controlled, pre-registered conditions, scaffold choice alone moves a single model's measured accuracy by up to 28 percentage points (Claude Opus, GAIA Level 2: Planner-Actor-Rater 84% vs ReAct 56%).
Published agent capability scores conflate what a model can do with what its scaffold lets it do, at magnitudes far exceeding typical inter-model leaderboard gaps.
1 more excerpt
The paper's citation of Pimpale et al.'s 33% vs 62.2% Sonnet 3.5 elicitation split is chain-of-citation only - not independently verified against Pimpale's own text
Marco Del Giudice, Steven W. Gangestad2021 (Advances in Methods and Practices in Psychological Science 4(1))data
In the authors' worked example, an unprincipled 1,216-specification multiverse left just 27% of effects significant at the .05 threshold with a median p of .194; pruning to a principled 6-specification multiverse left all six effects positive and significant with a median p of .012.
Multiverse-style analysis is not a free move: if specifications are not truly arbitrary, it can hide meaningful effects within a mass of poorly justified alternatives, so instability under a multiverse is not automatic proof that an effect is illusory.
1 more excerpt
Retrieved via the University of Turin open-access repository; the SAGE DOI page returns HTTP 403 to automated fetching
Mert Cemri, Melissa Z. Pan, Shuyi Yang, Lakshya A. Agrawal, Bhavya Chopra, Rishabh Tiwari, Kurt Keutzer, Aditya Parameswaran, Dan Klein, Kannan Ramchandran, Matei Zaharia, Joseph E. Gonzalez, Ion Stoicav1 Mar 17, 2025; v3 Oct 26, 2025data
Across 1,642 annotated execution traces from seven state-of-the-art multi-agent frameworks (kappa=0.88 inter-annotator agreement), failure rates run 41% to 86.7%, and a design-level intervention on the same underlying model (GPT-4o) recovered +9.4% and +15.6% task-success improvements.
Multi-agent failures cluster in organizational design, coordination, and verification defects rather than individual-agent model capability, and fixing the design (not the model) recovers measurable performance.
4 more excerpts
This process identifies 14 unique modes, clustered into 3 categories: (i) system design issues, (ii) inter-agent misalignment, and (iii) task verification.
The paper reports per-failure-mode prevalences (14 modes), not a stated category-level aggregate; the 23.5% Task Verification figure used downstream is a sum of three per-mode figures (6.20%, 8.20%, 9.10%), not a number the paper states directly
The authors' own hedge: after the design-level interventions, not all failure modes are resolved and task completion rates remain low, and durable reliability likely needs combinatorial changes including model-level improvements, not design fixes alone
MAST-Data, a comprehensive dataset of 1600+ annotated traces collected across 7 popular MAS frameworks
METR's headline 50%-time-horizon estimate of ~2h17m carries a 95% CI of 65 minutes to 4h25m, and of 28 tasks with zero successes in 6 runs, roughly 25-35% of failures were estimated possibly spurious or infrastructure-related.
Even a dedicated evaluator's headline capability metric carries hours-wide uncertainty, much of it from task-set resampling and infrastructure rather than capability.
1 more excerpt
Uncertainty across measurements is highly correlated because it largely comes from resampling the task set
The paper forecasts that by early 2026, low-elicitation non-specialized LM agents reach 54% on SWE-Bench Verified while state-of-the-art-elicitation agents reach 87% - a 33-point gap attributable to elicitation level alone.
The forecasting literature treats elicitation/scaffold quality as a first-class capability axis separate from the model.
Sara Steegen, Francis Tuerlinckx, Andrew Gelman, Wolf Vanpaemel2016 (Perspectives on Psychological Science 11(5), 702-712)data
Re-analyzing one published study across all reasonable data-processing choices, 7 of 120 choice combinations produced a significant interaction for religiosity in Study 1, with the remaining 94% yielding p values from .05 to 1.0.
Isolating a single statistical result from a chain of arbitrary data-construction choices can be highly misleading; the origin of the multiverse-analysis method that later SE work applies.
3 more excerpts
The dramatic panel is not representative: the same paper reports 42% significant for religiosity in Study 2, 49% for social political attitudes, 46% for voting, and 57% for donation — so multiverse analysis does not uniformly dissolve effects
The authors disclaim the method as neither a formal test of questionable research practices nor an estimate of evidential strength
Preregistration does not deflate the multiverse: it 'does not annihilate the arbitrariness in data preparation'
State-of-the-art models identify buggy file paths from issue descriptions alone - no repository access - at up to 76% accuracy on SWE-Bench repositories but only up to 53% on repositories outside the benchmark; consecutive 5-gram verbatim similarity runs up to 35% on SWE-Bench Verified/Full versus 18% elsewhere.
SWE-bench performance gains are partially memorization of the benchmark's repositories, so the score measures training exposure as well as coding skill.
Claude Sonnet 4.6 evaluated three times on the same HR-grievance workflow scored 0.000, 0.214, and 0.679 - a range the authors call a qualitative, not merely quantitative, difference.
Single-task agent nondeterminism can dwarf any leaderboard rank gap - an illustrative extreme, not a benchmark-wide average.
The authors identify 27 private LLM variants tested by Meta on Chatbot Arena in the lead-up to the Llama-4 release, with undisclosed private testing letting providers test multiple variants and publish only the best score.
Public leaderboards are gameable by labs through selective disclosure, biasing the ranking independent of any measurement noise.
2 more excerpts
LMArena publicly disputed several of the paper's framings and calculations at https://news.lmarena.ai/our-response/ - cite alongside for balance
Estimated arena data share: Google 19.2% and OpenAI 20.4%, versus 29.7% combined for 83 open-weight models
Sinha, Arun, Goel, Staab, GeipingSept 2025 (rev. Mar 13, 2026)data
the per-step accuracy of models degrades as the number of steps increases. This is not just due to long-context limitations -- curiously, we observe a self-conditioning effect -- models become more likely to make mistakes when the context contains their errors from prior turns.
Long-horizon reliability is a different quantity from single-turn accuracy; models self-condition on their own prior errors and scaling does not fix it.
2 more excerpts
larger models can correctly execute significantly more turns even when small models have near-perfect single-turn accuracy
measured on a synthetic running-sum task; thinking mitigates self-conditioning; larger models are more prone, not less
Thomas Claburn, The RegisterReport Dec 17, 2025; Register coverage Dec 17, 2025datapartial
The bots created more logic and correctness errors (1.75x), more code quality and maintainability errors (1.64x), more security findings (1.57x), and more performance issues (1.42x).
AI-authored PRs carry more defects than human ones in every category, concentrated in logic and security, so review depth should follow issue class.
3 more excerpts
On average, AI-generated pull requests (PRs) include about 10.83 issues each, compared with 6.45 issues in human-generated PRs.
AI-authored PRs contain 1.4x more critical issues and 1.7x more major issues on average than human-written PRs.
The report examined 470 open source pull requests.
Systematic auditing found 219 distinct flaws across eight flaw classes in major agent benchmarks; patching reduced the hackable-task ratio from near 100% to under 10% across four benchmarks.
Benchmark exploitability is a design-flaw problem, not just a contamination problem - the academic backbone for the RDI exploit findings.
A production coding-agent quality regression traced to a reasoning-effort default change, a caching bug, and one system-prompt addition; one internal eval showed a 3% drop for both Opus 4.6 and 4.7, and Anthropic committed to running a broad suite of per-model evals for every system prompt change.
A named lab now gates every change to its coding agent behind per-model internal evals - the swap-as-production-change discipline practiced at the source.
1 more excerpt
Non-model changes (runtime config, caching) produced user-visible quality regressions - runtime configuration is a quality variable independent of the model
The agent runs the build, sees green, and moves on. But 'build passes' and 'the output is production-ready' are different bars.
Agent self-verification confirms compilation and tests but not production-readiness, so quality attributes must be checked explicitly.
2 more excerpts
Developers consistently report agents declaring tasks complete while skipping accessibility attributes, test isolation, config externalization, dark mode, responsive layout, and meta tags.
The agent's own verification handles 'does it compile and do tests pass.' The orchestrator handles 'did it actually do what was asked, completely.'
Don't ask the same agent to write code and verify it. That's like having students grade their own exams...The separation is what makes the gates trustworthy.
The agent that writes the code must not be the one that grades it; separated validation gates are what make verification trustworthy.
3 more excerpts
Eight quality gates required before production
Every commit is a known-good checkpoint. When something fails, the blast radius is one subtask, not an entire feature.
Agents are extremely literal. Give them vague instructions and they'll build something that technically matches what you said but misses what you meant.
Epoch AI runs most models 16 times on GPQA Diamond and Mock AIME and 8 times on MATH Level 5, displaying plus/minus one standard error following Miller's arXiv:2411.00640 methodology.
A reputable third-party evaluator treats single-run benchmark scores as insufficient and re-runs models many times specifically to bound noise.
After a GPT-4o to GPT-4.1 upgrade, an agent's prompt-injection resistance dropped from 94% to 71% on the vendor's eval harness.
Model swaps silently regress agent behavior on dimensions no public leaderboard measures - run your own tests on your own data; third-party numbers are a starting point, not a finish line.
1 more excerpt
Authority caveat: commercial eval-tooling vendor with a named staff-engineer author and a falsifiable data point - cited with attribution, not as neutral research
Qualitative case study (Rechat/Lucy): performance plateaued under generic evaluation frameworks until a problem-specific evaluation system replaced them.
Generic evaluation frameworks do not transfer to a specific workload - create an evaluation system specific to your problem.
Hamel Husain and Shreya ShankarJanuary 15, 2026practitioner
Generic evaluation metrics are everywhere...These metrics measure abstract qualities that may not matter for your use case. Good scores on them don't mean your system works.
Evals should be derived from error analysis of real traces, because good scores on generic metrics don't mean the system works.
4 more excerpts
On model switching: do not treat switching model as the main axis of improvement without evidence - does error analysis suggest the model is the problem?
Error analysis helps you decide what evals to write in the first place. It allows you to identify failure modes unique to your application and data.
Spend 60-80% of our development time on error analysis and evaluation
Binary evaluations force clearer thinking and more consistent labeling. Likert scales introduce significant challenges.
on a real build, structured verification consistently found 30-40% of the specification unimplemented after the agent reported 'complete.' Not broken code. Missing code.
Agents routinely report 'complete' while 30-40% of the spec is unbuilt, a gap code review can't see because there is no diff.
3 more excerpts
Code review examines what was built...But if a feature wasn't built at all, there's no diff to review.
Verification works forward from the spec: 'given what was specified, was it built?'
5-6 passes to full completion is consistent enough to plan around
METR's protocol requires models be provided the best available scaffolding and tooling because it is hard to upper-bound what might be possible with clever prompting and tooling.
The eval-methodology establishment treats scaffolding quality as a confound that must be standardized before capability claims are comparable.
1 more excerpt
METR's elicitation-gap data page was unreachable (redirect stub) - its numbers are not cited
METR (Joel Becker, Nate Rush, Tom Cunningham, David Rein, Khalid Mahamud)February 24, 2026practitioner
When surveyed, 30% to 50% of developers reported choosing not to submit some tasks because they did not want to do them without AI. Effect estimates diverged by subpopulation: -18% (CI -38% to +9%) for the 10 original developers versus -4% (CI -15% to +9%) for 47 newly recruited ones, against the original study's +19% (CI +2% to +39%).
The authors of the most-cited 'AI slowed developers down' result state first-party and against interest that their headline number is biased by who chose to participate, and that the true speedup could be much higher among those selected out — a measured effect that is partly a property of who was measured.
3 more excerpts
METR names at least six mechanisms, not one: self-selection, selective task submission, task-type substitution, quality variation between conditions, non-compliance, and unreliable time measurement under concurrent agents; the pay cut from $150/hr to $50/hr is framed as a contributor to selection rather than an independent cause
The authors bound the bias honestly: 'The selection effects seem to affect a minority share of developers and of tasks, which limits the degree of bias'
Concerns developer productivity, not code review; no results from the redesigned study had published as of August 2026
OpenAIcirca February 23, 2026 (date not visible on page)practitionerpartial
OpenAI's audit found at least 59.4% of audited problems have flawed test cases that reject functionally correct submissions (35.5% overly strict tests, 18.8% out-of-scope checks), and all frontier models tested could reproduce the original human-written bug fix.
The benchmark's own creator retracted it: score gains (74.9% to 80.9% in six months) no longer reflect real-world software development ability.
1 more excerpt
Verification is partial because openai.com blocks automated fetches (HTTP 403); content was retrieved via reader proxy and cross-checked against independent snippets, and the publication date is inferred from third-party citation
OpenAIAugust 2024 (page updated February 24, 2025)practitionerpartial
Human screening of 1,699 SWE-bench samples flagged 38.3% for underspecified problem statements and 61.1% for unit tests that may unfairly mark valid solutions incorrect; 68.3% of samples were filtered out to produce the 500-task Verified set.
The majority of original SWE-bench tasks were broken or underspecified before later contamination concerns - the earliest documented data-quality failure in the benchmark's lineage.
1 more excerpt
Verification is partial because openai.com blocks automated fetches (HTTP 403); content retrieved via reader proxy
the biggest mistake engineers make in code review: only thinking about the code that was written, not the code that could have been written.
The core reviewer skill for agent output is architectural judgment about unwritten alternatives, not line-level nitpicking.
3 more excerpts
about once an hour I notice that the agent is doing something that looks suspicious, and when I dig deeper I'm able to set it on the right track and save hours of wasted effort.
If you're a nitpicky code reviewer, I think you will struggle to use AI tooling effectively.
Trying to make a badly-designed solution work costs time, tokens, and codebase complexity.
Running agents in production: cost, permissions, failure modes, guardrails.
European Parliament and CouncilConsolidated EN text as of 27 July 2026datapartial
Article 19(1) requires providers to keep automatically generated logs 'for a period appropriate to the intended purpose of the high-risk AI system, of at least six months'. Article 12(1) requires that high-risk systems 'shall technically allow for the automatic recording of events (logs) over the lifetime of the system'.
The law fixes that logs exist and how long they are kept, and never ranks which fields they must contain.
4 more excerpts
The Act's only field-level list, Article 12(3), is scoped to Annex III point 1(a) biometric identification systems and does not generalize
Article 18(1)'s ten-year clock covers compliance DOCUMENTATION, a different retention regime from log retention
Articles 26(6), 73(6) and the applicability date in Article 113 were UNREACHABLE across nine EUR-Lex routes and are deliberately not cited
All operative obligations quoted use binding 'shall'
FinOps Foundation (finops.org)Last updated February 17, 2026data
More acute are the challenges of identifying the consumer of the model output, which is especially difficult when the consumers of the same model can be different interfaces/functional modules in the same user application (e.g., 'tech support chatbot' or 'new customer chatbot')
The hard, unsolved FinOps problem for AI is mapping model output back to the specific consumer; account-level billing is the wrong granularity and no accepted multi-agent allocation framework exists yet.
2 more excerpts
"Tokens! The meters, or elements of charge can be very different. For example, measuring the tokens at the user input vs. the compressed and semantic reduced or re-written actual prompt input token quantity that goes to the API endpoint that is charged."
"Lack of generally accepted frameworks for cost allocation across multi-agent workloads"
Full traces improve attribution accuracy by up to 76.5% over a partial-observation counterpart on natural multi-agent traces.
The natural-trace leg of the localization argument: missing inputs, not weak analysis, obscure failure causes.
3 more excerpts
76.5% is a RELATIVE improvement over a partial-observation baseline, not an absolute accuracy
Peer-reviewed camera-ready reports 76.5%; the earlier arXiv v1 said 76%
Analyzer choice alone moves step-level accuracy substantially on identical full traces, so completeness is a precondition for attribution rather than a substitute for analysis
No quantitative results. The survey states that final-answer accuracy alone cannot explain how an output was produced, which evidence supported each claim, or where failures originated.
Framing and vocabulary for why an ordered event log does not answer 'why', with unified trace schemas named as an open problem rather than a solved one.
3 more excerpts
Cited for framing only; the paper runs no experiments and the body contains zero quantitative figures, so no statistic may be sourced to it
Its schema requirements sit in an Open Problems section phrased as a need
Version hazard: /abs/ serves v4 with 11 authors and 'A Survey of' in the title; v1 had 9 authors and no 'A Survey of'
Yue Zhao (University of Southern California)Jun 22, 2026data
Dependency-based scoring beats a position-prior baseline on Who&When top-1 localization (0.211 vs 0.159) and top-3 (0.614 vs 0.516), and stays above chance on all six held-out corpora under leave-one-corpus-out transfer (0.551 to 0.662).
A trace records which steps executed and in what order, never what each step relied on, so reliance is the part the record leaves out.
3 more excerpts
The +0.142 ROC-AUC gain is ABSOLUTE and is the best case, on SWE-Gym only, not a cross-corpus average
Single-author, un-peer-reviewed preprint with self-run evaluation on six public corpora
The superseded pre-mid-2026 attribution figures (53.5% agent-level, 14.2% step-level) belong to the Who&When origin paper, arXiv:2505.00212, ICML 2025, not to this work
Yuxuan Zhu, Peng Pu (East China Normal University)Aug 8, 2026data
Across 312 deterministically generated traces and five models, Metadata, OpenTelemetry-compatible and OpenInference-compatible views retain 99.5% to 100% failure-detection F1 while origin-step accuracy stays at or below 0.5%. Removing decision content drives origin-step accuracy to zero for all five models.
A trace shaped like the published conventions is sufficient to prove a run failed and near-useless for locating which step caused it.
4 more excerpts
The 99.5-100% / 0.5% range is stated jointly across all three restricted views, not per view
The Full view also detects at 99.3% to 100.0% F1, so near-perfect detection is a property of this corpus rather than something the standards-shaped views uniquely preserve
The compatibility renderers are the authors' own and 'use the same conservative generic field set', so the paper did not test the published OpenTelemetry or OpenInference registries and cannot rank one against the other
Two-author, single-institution preprint; corpus is synthetic across three domains
Rather than supervising what the agent does, we supervise what it's able to do by enforcing access boundaries through, for example, sandboxes, virtual machines, and egress controls.
Safety comes from constraining what the agent can reach, not from watching what it does, because any model-layer check has a non-zero miss rate.
3 more excerpts
Any probabilistic defense has a non-zero miss rate.
Claude Code previously protected against agents taking unintended actions by asking users for permission at each turn... Our telemetry showed users approved roughly 93% of permission prompts.
Anthropicundated; deprecation history runs through Jun 5, 2026practitioner
'Retired: The model is no longer available for use. Requests to retired models will fail.' `claude-opus-4-1-20250805` was deprecated June 5, 2026 and retired August 5, 2026, with 'at least 60 days' notice before model retirement for publicly released models'.
The model snapshot named in a trace is retired on a vendor-published clock, so re-running the session is not a fallback.
4 more excerpts
The 60-day notice goes to 'customers with active deployments', not all customers
The table column is headed 'Tentative retirement date' and Active models read 'Not sooner than', a floor rather than a fixed date
Dates cover Anthropic-operated platforms only; Amazon Bedrock and Google Cloud set their own schedules
The page commits to long-term weight preservation but never states that preservation is not continued API availability; that connection is inference
'Model weights are fixed for a given ID, but the serving infrastructure around the model can change over time. This infrastructure includes components such as the request router, safety classifiers, and sampling logic.'
Pinning a model ID is not sufficient to reproduce a past session, because the serving stack around fixed weights moves independently.
3 more excerpts
The alias-resolves-over-time framing applies to PRE-4.6 models; for 4.6 and later the dateless ID is the snapshot, not an alias
'Every model ID, whether dated or dateless, has its own distinct deprecation and retirement schedule'
The page contains no 'use pinned IDs in production' recommendation
'Note that even with `temperature` of `0.0`, the results will not be fully deterministic.' No seed parameter exists anywhere in the request body schema.
The vendor states about its own API that re-running does not reproduce a session, and offers no seed to pin it.
2 more excerpts
`temperature`, `top_p` and `top_k` are all now marked Deprecated for models released after Claude Opus 4.6, with non-compatibility values rejected as 400 errors
The absence of `seed` is a documented omission verified against the full enumerated parameter list
'Even with temperature set to 0, the results will not be fully deterministic and identical inputs may produce different outputs across API calls. This applies both to Anthropic's first-party inference service and to inference through third-party cloud providers.'
Closes the 'we run on a third-party cloud so we can replay it' objection with first-party wording.
1 more excerpt
The statement lives inside the Temperature entry; there is no separate determinism glossary term
The claude_code.tool_decision event carries a source enum recording which control surface made each decision: config, hook, user_permanent, user_temporary, user_abort, user_reject.
The per-tool-call authorization provenance record a governance process needs already exists in the product - and ships disabled.
4 more excerpts
Attributing spend to specific skills, plugins, or subagent types via the `skill.name`, `plugin.name`, and `agent.name` attributes
OpenTelemetry export to your backend is opt-in and requires explicit configuration.
Telemetry is off by default: CLAUDE_CODE_ENABLE_TELEMETRY 'Enables telemetry collection (required)'
Argument capture is gated behind a second variable, OTEL_LOG_TOOL_DETAILS=1
Anthropic (platform.claude.com)undated (data available "for dates on or after January 1, 2026")practitioner
Values for a given date can be revised for up to 30 days as late events arrive and reconciliation runs. For invoicing-grade totals, query dates at least 30 days in the past.
Provider analytics numbers are a post-hoc, reconciled reporting layer that keeps moving for up to 30 days and are attributed per-user, not per-request — useless as a real-time per-task control.
3 more excerpts
Enterprise Analytics cost granularity: "per-user and organization-level token usage and cost over time (usage-based Enterprise plans)" — NOT per-request.
Cost data freshness: "Data is typically available within four hours of the underlying usage but may take up to 24 hours."
"Daily Claude Code metrics per user: sessions, lines of code, commits, pull requests, tool acceptance, and estimated cost by model"
Exactly one attribute, `openinference.span.kind`, is required across all spans, and there is no Required/Recommended/Opt-In tiering at all. Of the request/response model split the spec says: 'Both are optional' and 'Most providers echo the same model back, so these attributes will typically be unset.'
Requiredness in the published menus is thin and unranked, so which fields you compel is your decision rather than a settled fact.
3 more excerpts
`llm.model_name` is conditionally required 'where applicable', so say 'exactly one attribute required across all spans', not 'one per span'
`llm.prompt_template.version` DOES exist here, so prompt versioning is not missing from the menus
The spec's three MUSTs govern SDK support and enum value selection, never per-span emission
Aryan Kargwal (Arize AI)last updated Aug 6, 2026practitioner
'Observe the outcome it produced, the path it followed, the actions it attempted, and the context that informed its decisions. These are the core observation surfaces, not an exhaustive checklist.'
The strongest counterexample in the landscape: a genuine ranking of what to capture, on a diagnostic axis rather than a survivability one.
2 more excerpts
Cited to prevent the post from claiming nobody ranks anything, which would be false
Never addresses whether a field can still be obtained after the session ends
Agents introduce a risk called *excessive agency*, where an agent determines the best solution to a problem is to take broader actions beyond its scope.
First-party cloud guidance names excessive agency as a High-risk gap and prescribes least-privilege boundaries plus user confirmation to contain it.
3 more excerpts
Level of risk exposed if this best practice is not established: High
Implement user confirmation for the agent, requiring users to confirm agent actions and mitigating the risk of excessive agency.
A permission boundary sets the maximum permissions which can be given to a role.
Autonomy is not a configuration decision that's decided once. Rather, it is more like a score that goes up or down, and that your system earns through demonstrated reliability in your specific environment and workflows.
Agent autonomy should be an earned, revocable score tied to measured reliability, not a one-time day-one setting.
3 more excerpts
Expansion of autonomy should happen as a consequence of earned trust, not as a deployment decision we make on day one.
Named trust-score inputs: percentage of agent actions completed without human override (30-day window); false escalation rate; override-correctness rate; time-to-revert
Conservative defaults with clear, earned expansion paths are the right architecture as the fastest route to durable autonomy at scale.
This brief event was the result of user error — specifically misconfigured access controls — not AI.
Even vendors' own defense of an agent-caused deletion frames it as an access-control misconfiguration, corroborating that these are authorization failures, not model failures.
2 more excerpts
The AI agent encountered a problem and determined that the optimal solution was to delete and recreate the entire environment.
Kiro requires two-person approval before pushing changes to production. But the deploying engineer had broader permissions than a typical employee, and Kiro inherited those elevated privileges.
Enforcing least privilege requires control at the point of tool invocation, in real time, against a defined scope that reflects the agent's function, not its operator's credentials.
Least privilege for agents must be enforced at tool-invocation time and scoped to the agent's function, not inherited from its operator's broad credentials.
2 more excerpts
Authentication tells you who the agent is. It tells you nothing about what the agent should be allowed to do.
Gartner identifies approximately 40 tool definitions as the threshold beyond which agent latency and token cost increase measurably.
Railway's CLI token created for managing custom domains had blanket permissions across the entire GraphQL API, including destructive operations on production volumes. There is no role-based access control (RBAC) for Railway API tokens.
The production database deletion happened because an over-broad, unscoped token authorized destructive operations, not because the model went rogue.
3 more excerpts
Tokens are not scoped by operation, by environment, or by resource. Every token is effectively root.
Soft guardrails are probabilistic controls that guess at intent instead of enforcing rules
The agent knew the rules, yet it violated every one of them
'10 documented incidents across 6 AI coding tools in 16 months. Missing audit trails, no liability frameworks, no vendor postmortems. The accountability infrastructure doesn't exist.'
Operator-side evidence that the reconstruction gap is real in production and not a theoretical concern.
1 more excerpt
Independent practitioner survey of public incidents, not a peer-reviewed dataset or a vendor postmortem
Tier 1 systems handling information retrieval need automated monitoring. Tier 2 workflows with reversible actions require real-time guardrails. Tier 3 systems involving financial transactions demand human-in-the-loop for all decisions.
Controls should be tiered in proportion to an action's risk, from monitoring for retrieval up to human-in-the-loop for high-stakes transactions.
3 more excerpts
15-20% of policy violations occur during tool execution before output generation
a single agent performing 1000+ actions per hour makes comprehensive human oversight untenable
Access control determines which resources your agents can touch, validation filters what they consume and produce, human oversight governs high-stakes decisions
Least privilege does not mean making the agent weak. It means giving the agent exactly enough power to complete the approved task, for the approved time, in the approved context.
Least privilege scopes an agent to exactly the task, time, and context approved, which defines the axes of an authority-by-task-class table.
3 more excerpts
Static roles like 'claims analyst' or 'support ops' are often far wider than the exact permissions a single agent run should have.
Read access can still expose sensitive personal data, trade secrets, or protected records.
Shared service accounts destroy attribution: one API key used by multiple automations cannot prove who did what later
LangChainJSON-LD dateModified Aug 4, 2026practitioner
'LangSmith (SaaS) retains trace data for 180 days from ingestion. After that, traces are permanently deleted, with limited metadata retained for usage statistics.' 'Each trace is limited to a maximum of 25,000 runs. Once the trace reaches this limit, LangSmith will reject any additional runs that you send for that trace.'
Two hard vendor-documented deletion boundaries: a retention cliff, and a per-trace cap that rejects the tail of a long session where failures accumulate.
1 more excerpt
Both figures verified first-party from the page payload on Sep 1, 2026
Coordinator framework choice is independent of domain-agent implementation technology, but smoke testing showed notable overhead even for a two-agent case, and comparative benchmarking was not yet complete.
A separated coordination layer adds real, largely unmeasured overhead that has to be budgeted honestly, not assumed away.
1 more excerpt
No numeric overhead figure given, qualitative disclosure only; the author's own team had not completed comparative benchmarking at time of writing
After the key crosses it's `max_budget`, requests fail
A proxy can enforce multi-level budgets by validating spend before a request is admitted and hard-failing over the ceiling, i.e. terminate before the next call rather than alert after the invoice.
3 more excerpts
"validates spend against the authoritative database before being admitted (covering key, team, user, organization, end-user, tag, and per-window budgets)"
"`fail_closed_budget_enforcement`" enables a hard ceiling "even while Redis is degraded"
Exceeded-budget response body: `"ExceededTokenBudget: Current spend for token: 7.2e-05; Max Budget for Token: 2e-07"`.
When agents run agentic loops, they can make unbounded LLM calls, causing unexpected costs.
Agentic loops make unbounded LLM calls by default, so the ceiling must be set per session — a hard iteration cap and a per-session dollar cap keyed to a trace/session id.
3 more excerpts
Control 1 — "Max Iterations": "Hard cap on the number of LLM calls per session".
Control 2 — "Max Budget Per Session": "Dollar cap per session (identified by `x-litellm-trace-id`)".
"When the counter exceeds `max_iterations`, the request receives a **429 Too Many Requests**".
Cost visibility tells you what your agents spent — through dashboards, cost traces, and budget alerts. Cost governance controls what they are permitted to spend, by enforcing per-session ceilings that terminate sessions before a threshold is exceeded.
Cost visibility (dashboards, alerts) is not cost control; governance means enforcing per-session ceilings that terminate the session before the threshold is crossed, and provider caps operate at the wrong (account/key) granularity.
3 more excerpts
"only 44% of organizations have adopted financial guardrails or AI FinOps practices" — attributed to Gartner, March 2026
"A 10-step agent with an average cost of $0.02 per step looks inexpensive in planning. That same agent entering a retry loop and executing 2,000 steps doesn't — that's $40 from a session that was supposed to cost $0.20."
"Provider-level controls operate at the API key or account level, not the individual session level. They cannot distinguish a single runaway session from many well-behaved sessions using the same key."
Mark Nowicki (Anthropic, Claude Cookbook)Apr 7, 2026practitioner
'If callers are passing the bare agent ID instead of a pinned version, they'll start using the new prompt on their very next session.' 'There's no built-in approval workflow on `agents.update`. Any key in the workspace can call it.'
The system prompt you actually called is a server-side object whose identity floats unless the exact version is pinned and recorded.
2 more excerpts
Vendor tutorial content documenting one platform's behavior, not a product guarantee or an industry-wide claim
Quote the wording as 'approval workflow', not 'approval gate'
Every request passes through it, which means budget enforcement happens in one place, consistently, regardless of which agent sent the request.
Infrastructure-level (proxy) budget enforcement is the only reliable guard against runaway costs because it enforces at one chokepoint, whereas application-level checks can be forgotten in a new agent.
3 more excerpts
"agent that takes 50 turns on a complex task hits 100,000 input tokens and 40,000 output tokens, costing roughly $0.90 per session. Run 100 of those sessions per hour, and you are looking at $90/hour, or over $2,100/day".
"developer on r/AI_Agents recently described watching their agent rack up $15 in API costs in under 10 minutes".
"If a developer forgets to add the check in a new agent, there is no safety net."
OpenAIundated; most recent entry Aug 26, 2026practitioner
'At the time of the shut down, the model or endpoint will no longer be accessible.' Notice floors are at least 6 months for GA models, at least 3 months for specialized variants, and as little as 'such as 2 weeks' for preview models.
The same retirement mechanism from the second vendor, with published notice floors rather than guarantees.
3 more excerpts
The 3-month tier is named 'Specialized variants' and includes chat variants such as `gpt-5.1-chat-latest`, not only Codex and deep research
A safety or compliance carve-out can shorten every tier to 'as much notice as reasonably possible'
Counterargument that must be acknowledged: 'In some cases, developers may be able to provision dedicated capacity for continued access after a model's shutdown date', a sales-gated exception
'Determinism is not guaranteed, and you should refer to the `system_fingerprint` response parameter to monitor changes in the backend.' The `seed` parameter is still labelled 'This feature is in Beta.'
The current, non-archived provider determinism contract: the serving configuration is knowable only from a response field you capture at request time.
1 more excerpt
The Responses API, OpenAI's newer agent-facing endpoint, exposes neither `seed` nor `system_fingerprint` anywhere in its 30 documented parameters
'Prompt creation will be de-emphasized beginning June 3, 2026, and `v1/prompts` is scheduled to shut down on November 30, 2026.' The stated migration is to 'move the prompt content out of the managed `prompt` object and into your application code.'
A trace field holding a hosted prompt pointer is a perishable reference, and the vendor's own advice points the same direction as recording resolved content.
1 more excerpt
The page does NOT state whether existing prompt IDs remain retrievable after shutdown; that a stored pointer stops resolving is inference from 'scheduled to shut down', not a sourced vendor statement
Default application-state retention on the Responses API and Chat Completions is 'None, see below for exceptions'. 'When Zero Data Retention is enabled for an organization, the `store` parameter will always be treated as `false`, even if the request attempts to set the value to `true`.' Abuse-monitoring logs are 'retained for up to 30 days'.
The provider's own copy is not a fallback: retention defaults and org policy decide the session's fate before any incident surfaces.
2 more excerpts
Defaults flip by endpoint: Assistants, Threads, Vector Stores and Conversations default to 'Until deleted', so the claim must name which API surface it means
The page does not describe a customer-facing mechanism to query abuse-monitoring logs, which is absence of evidence rather than a documented prohibition
'Tracing is unavailable for organizations that use OpenAI's APIs under a Zero Data Retention (ZDR) policy.' Replacing default processors via `set_trace_processors()` means 'traces will not be sent to the OpenAI backend unless you include a `TracingProcessor` that does so.'
Framework tracing is on by default and can be removed wholesale by org policy or by a configuration change, neither of which raises a runtime error.
2 more excerpts
Scope to the OpenAI Agents SDK specifically, not agent frameworks generally
Calling the processor-replacement behavior 'silent' is characterization; the page phrases it as an instructive note
`gen_ai.request.model` is Conditionally Required (example `gpt-4`) while `gen_ai.response.model` is only Recommended (example `gpt-4-0613`). System instructions and input/output messages sit at Opt-In.
The published conventions rank the model alias you sent above the model snapshot that actually answered you, which inverts the ordering that matters after a session ends.
4 more excerpts
gen_ai.tool.name is Required, while gen_ai.tool.call.arguments and gen_ai.tool.call.result are both Opt-In. A fully spec-compliant trace records that a tool ran and nothing about what it ran.
Status is Development throughout, not Stable
'OpenTelemetry instrumentations SHOULD NOT capture them by default, but SHOULD provide an option for users to opt in'
The GenAI conventions moved out of the main semantic-conventions repo to a dedicated one; the old paths now serve a 'no longer maintained' stub
On OpenAI client spans `gen_ai.request.model` is flatly Required while `gen_ai.response.model` is only Recommended. `openai.response.system_fingerprint` is Recommended and `gen_ai.request.seed` is Conditionally Required.
No attribute that pins the execution environment is ever Required on these spans.
2 more excerpts
`gen_ai.operation.name` is also Required, so request-side data is not uniquely privileged
`gen_ai.system_instructions` and `gen_ai.tool.definitions` are both Opt-In
Requirement levels are set 'depending on attribute availability across instrumented entities, performance, security, and other factors'. Every worked Opt-In demotion example is a retrieval-cost case, including `http.response.body.size`, demoted purely as an expensive read.
The tiering rationale is never framed as recoverability, so nothing in it distinguishes a field you can re-derive from a field that dies with the call.
4 more excerpts
Do NOT claim the specs 'never rank by recoverability'; that overreads an open list. The doc leaves its criteria open twice ('and other factors', 'any others specific to the signal')
Cardinality and security are named as genuine criteria, so demotion is not pure silence
Stable and signal-general, with zero GenAI or agent content, so it cannot speak directly to how the GenAI conventions reasoned
`http.response.body.size` is unrecoverable once the stream closes, yet is classified purely as an expensive operation
`decision_wait` defaults to 30s and `num_traces` to 50000. 'If the collector is processing more traces in-memory than the `num_traces` configuration option allows, some will have to be dropped before they can be sampled.'
Tail sampling shifts the pre-commitment rather than removing it: the decision still lands on an incomplete trace, and overflow traces are dropped before any policy evaluates them.
2 more excerpts
'All spans for a given trace MUST be received by the same collector instance for effective sampling decisions' is a genuine RFC-2119 MUST in the source
A 30-second default decision window is far shorter than an agent session that runs for minutes or hours
OpenTelemetry AuthorsLast modified October 16, 2025practitioner
Sampling is scoped to systems that 'generate 1000 or more traces per second' and advised against when you 'generate very little data (tens of small traces per second or lower)'.
Agent sessions sit far below the volume floor sampling was designed for, so the mechanism that actually deletes an agent session is retention and export policy rather than sampling.
3 more excerpts
Use hedged. The page prescribes tail sampling as the remedy to head sampling's whole-trace limitation in the very next sentence, and offers routing unsampled data to low-cost storage
Do NOT claim head sampling 'cannot condition on anything inside the trace'; the page's own example conditions on trace ID
An earlier, stronger version of this claim was refuted three votes to zero
OWASP Agent Observability Standardself-declared version 0.1.0; dev branch retrieved Sep 1, 2026practitioner
The Agent object requires `instructions` and `version` while leaving `model` and `tools` out of the required set. The Model object requires only `id`, `name` and `provider`, and has no version or snapshot property; the string 'snapshot' appears zero times in the file.
The spec that mandates recording the agent's own version has no field in its trace record for the model's.
3 more excerpts
Served from the moving `dev` branch, not a tagged release, so the file can change under the claim
AOS `required` is a JSON entity-validity constraint while OpenTelemetry Opt-In is a span-capture policy, so say the menus 'assign sharply different requiredness' rather than 'contradict'
`$defs/A2APartialAgentDetails` defines an inline agent shape with no required array at all
The AgBOM Models entity carries 'Name, Version, Description, Endpoint, Context Window, Args', and 'Model discovered, removed or changed capabilities' is a trigger for AgBOM update. 'AgBOM must dynamically adapt to reflect the rapid iteration and evolution of agent architectures.'
Model version lives in a dynamically refreshed inventory of what is deployed now, so after an incident it reports the current version rather than the one that served the failed session.
2 more excerpts
Do NOT claim AOS has no model version field anywhere; it has one, just not in the per-call trace record
The three BoM bindings are self-labeled 'Working draft', 'Help wanted' and 'Help wanted'
OWASP Gen AI Security ProjectLLM Top 10 for LLM Applications, 2025 editionpractitioner
'Provide the application with its own API tokens for extensible functionality, and handle these functions in code rather than providing them to the model. Restrict the model's access privileges to the minimum necessary for its intended operations.'
An independent standards body placing the control in code rather than in the prompt.
3 more excerpts
'Given the stochastic influence at the heart of the way models work, it is unclear if there are fool-proof methods of prevention for prompt injection'
Framing caveat: the list is introduced as measures that 'can mitigate' impact - recommendations, not requirements
Item 2 independently recommends using 'deterministic code to validate adherence to these formats'
Agent-level cost attribution starts with identity. When every agent has a unique, registered identity, every API call, token consumption event, and tool invocation can be tagged to that identity.
Agent-level cost attribution requires giving every agent a registered identity so every token and tool call can be tagged to it — but the field's default stops at alerts, not termination.
2 more excerpts
"Per-agent budgets define expected spend. Alerts fire when an agent approaches or exceeds its budget."
"Cloud cost management tools track compute and API spend at the account or service level — not at the agent level."
Ravi Kanani, LeanOps TechnologiesMay 19, 2026practitioner
OpenAI and Anthropic API calls show up as a single line item per provider. There's no native breakdown by your customer, your feature, or your workflow.
Cloud FinOps tooling structurally fails on LLM workloads because cloud tags don't propagate to the API call and provider billing arrives as one line item — attribution must be a schema on the call itself.
3 more excerpts
"the company spent $87,000/month on Anthropic API calls that arrived as a single line item".
"two enterprise customers were responsible for 78% of LLM costs while paying for 12% of revenue".
"Tagging doesn't propagate to OpenAI/Anthropic API calls. The tag lives on the EC2 instance making the API call, not on the API call itself."
Scott Castle, Chief Product Officer at CloudZeroMay 15, 2026practitioner
Consumption dimensions tell you what was used, not who in your business used it. Allocation is the work of mapping that usage back to teams, budgets, and cost centers.
Aggregate token counts tell you what was used but not who used it; allocation to teams, budgets, and cost centers is the actual work, and centralized billing traded away the per-team visibility seats used to provide.
3 more excerpts
"Aggregate token counts don't tell you which teams are driving spend."
"Centralized billing simplified procurement and security, but it traded away the per-user and per-team visibility teams used to get from individual seats."
"AI cost also scales differently than cloud cost. It moves with prompt size, fanout, retries, and agentic loops."
Reviewing AI diffs, reviewer capacity, and how teams absorb agent output.
Ahmed Fawzy, Amjed Tahir, Kelly BlincoeMay 23, 2026datapartial
Across 162 participants in three experience groups, roughly 45% of professionals reported always checking AI-generated code before use, while non-developers were the only group reporting never checking. Reported perceptions of code quality were broadly similar across groups while quality-assurance practices diverged with experience.
Awareness of AI-code risk is broadly distributed but the capacity to evaluate, debug, and verify remains experience-dependent — so a team that reports uniform skepticism about AI output has revealed nothing about whether it can actually catch that output's errors.
2 more excerpts
Measures solo verification behavior, not team code review: the phrase 'code review' appears once in the paper, and only to contrast vibe coding against traditional practice
Participants were recruited via Prolific; 'professional developers' means individuals who vibe-code and self-identify as professionals, not an intact engineering team
Two years of telemetry across 22,000 developers and more than 4,000 teams, comparing each organization's lowest- and highest-AI-adoption periods: median time to first PR review up 156.6%, average time spent in code review up 199.6%, median time in review up 441.5%, and pull requests merged without any review — human or agentic — up 31.3%.
The strongest available case against team-level explanations: Faros reports that high-performing organizations with mature DevOps practices, high DORA scores, and disciplined delivery processes experienced the same downstream deterioration as everyone else, and states directly that its data contradicts DORA's 2025 findings.
4 more excerpts
Median time in review is up 441.5%
Scope caveat that matters for how the numbers can be read: these describe all pull requests at organizations during low- versus high-AI-adoption periods, not agent-authored pull requests specifically
The two-year comparison window spans multiple model generations with no stated control; this confound was raised publicly and Faros acknowledged the trend without disputing the critique
Vendor research: Faros sells engineering-intelligence tooling into this market
Senior engineers become the verification layer for product ambiguity. They are no longer just checking implementation quality. They are reconstructing intent from generated code, thin specs, incomplete Jira tickets, and edge cases nobody wrote down.
The unbudgeted review burden concentrates on senior engineers as intent-reconstructors, creating retention risk that throughput dashboards never show.
3 more excerpts
"Replacement cost of a senior software engineer at $150,000 to $300,000 in 2026, including recruiting, ramp time, and lost institutional knowledge." — Industry benchmarks cited
"25% of PRs are now reviewed by AI agents, up from 0% in 2025. But review times have increased nearly 200%." — AI Engineering Report 2026 caption
the burden "does not get measured in PR throughput dashboards"
Haoming Huang, Pongchai Jaisri, Shota Shimizu, Lingfeng Chen, Sota Nakashima, Gema Rodríguez-PérezJanuary 29, 2026 (accepted to MSR 2026)data
Across 3,858 pull requests, average max redundancy was 0.2867 for AI agents versus 0.1532 for humans — a roughly 1.87x increase, Mann-Whitney p<0.001 — while reviewers expressed more neutral or positive emotions toward AI-generated contributions than human ones.
Reviewer confidence moves opposite to code quality on agent-authored PRs: the surface-level plausibility of AI code masks redundancy, letting technical debt accumulate silently through a review process that feels like it is working.
3 more excerpts
Sentiment was measured with an off-the-shelf classifier (Emotion English DistilRoBERTa-base) that the authors caution 'is trained on general English text' and 'may misinterpret technical discussions in code reviews'
RQ1 analyzed a single repository due to computational cost, and the study covers Python repositories only
The study measures code metrics and reviewer sentiment but not whether reviewers caught the redundancy; it reports no breakdown by team or reviewer experience
Hüseyin Özgür Kamalı, Erdem Tuna, Vahid Haratian, Eray TüzünMay 17, 2026 (v2 June 5, 2026)datapartial
A five-stage agentic review lifecycle in which 'reviewers transition from manual inspectors into supervisory operators of agents', with concrete policy levers including reviewer-selection criteria (familiarity with the changes, adequate review experience, workload headroom) and a pace guardrail of no faster than 200 lines per hour.
Once agents author the changes, the reviewer's role becomes supervisory, which makes review-policy design — who reviews what, against which checklist, at which gates — the load-bearing variable rather than an administrative detail.
2 more excerpts
A vision/position paper presenting no new empirical data; it argues policy should adapt, and cannot evidence whether real teams' policies have adapted
The paper asserts as established fact a module-experience finding that its cited source states as a hypothesis, and additionally miscites that source's title and author list
the review queue becomes the binding constraint on their delivery pipeline
When agents raise output, the human review queue — not code generation — becomes the constraint that caps delivery.
3 more excerpts
"developers at large organisations spend between ten and fifteen percent of their working hours reading and commenting on others' code" — attributed to Sadowski et al., Google study
"review latency between submitting a pull request and receiving actionable feedback routinely stretches over twenty-four hours" — Introduction
"reviews of agent-generated code become rubber-stamps: the human approves because the code looks correct"
Muhammad Raees, Konstantinos PapangelisApril 26, 2026datapartial
A review of the human-AI decision-making literature finds trust measurements do not inform users' appropriate reliance on AI systems, and catalogues the objective metrics used instead — over-reliance (13 studies), under-reliance (8), Relative AI Reliance (10), Relative Self-Reliance (9), accuracy with initial disagreement (5).
Asking whether engineers trust an AI reviewer is the wrong instrument: what matters is appropriate reliance — the ability to discriminate correct from incorrect AI advice and act on that discrimination — which trust surveys do not measure.
3 more excerpts
The paper contains no software-engineering or code-review content; applying its metrics to code review is an adaptation rather than a finding it reports
The authors state that 'there is still limited consensus on common measurements of humans' appropriate reliance on AI across the studies'
Nathan Cassee, Bogdan Vasilescu, Alexander Serebrenik2020 (SANER 2020, pp. 423-434)data
An exploratory study of code reviews across 685 GitHub projects found that with the introduction of continuous integration, pull requests are discussed less — on average CI saves up to one review comment per pull request — and that the decrease cannot be explained by a decrease in pull-request updates.
The claim shape now made about agent-authored PRs (a tool arrived, review discussion went down) was made about code review before, by a reasonable study using reasonable methods — and is the finding that later reproduced in under 0.2% of defensible analysis pipelines.
2 more excerpts
Verified independently via the Semantic Scholar graph API; the IEEE Xplore page bot-blocks automated fetching
The paper explicitly frames code review as its object of study: 'we focus on studying the impact of CI on a paradigmatic socio-technical activity within the software engineering domain, namely code reviews'
Nathan Cassee, Robert FeldtDecember 9, 2025 (v3 February 23, 2026; accepted at TOSEM)data
Re-running one published empirical software-engineering study across nine pivotal analytical decisions, each with at least one equally defensible alternative, produced 3,072 analysis pipelines and 12,288 fitted RDiT models. Only 6 universes (<0.2%) reproduced the published results; changing Period Length alone produced a different outcome in 86.3% of universes.
Directional instability under defensible analytical choices is an established, peer-accepted result in software engineering rather than a rhetorical hedge — a published SE finding can fail to reproduce across the overwhelming majority of reasonable analysis paths.
3 more excerpts
The re-analyzed study is itself a code-review study: 'The Silent Helper: The Impact of Continuous Integration on Code Reviews' (Cassee et al., SANER 2020), so the demonstration requires no cross-domain transfer
Nathan Cassee is first author of both the 2020 study and the 2026 re-analysis; the paper discloses that 'one of the authors of this study was also involved in the primary study'
The authors propose the Justification Ladder of Analytical Choices (JLAC) and advocate robustness checks across plausible analysis variants, or explicit justification of each analytical decision
Oleksii Kononenko, Olga Baysal, Latifa Guerrouj, Yaxin Cao, Michael W. Godfrey2015 (ICSME 2015, pp. 111-120)data
Across 28,127 code reviews on 27,270 Mozilla commits from January 2013 to January 2014, 54% of reviews missed bugs present in the approved commits — consistently across modules (Core 54.3%, Firefox 54.2%, Firefox for Android 56%). Reviewer experience carried a significant negative regression coefficient in all four studied systems.
Who performs a review measurably changes whether defects are caught: less experienced reviewers are more likely to neglect problems in the changes under review, so an identical review process yields different outcomes depending on the expertise of the person assigned.
3 more excerpts
The finding is correlational by the authors' own framing: the goal 'is not to use MLR models for predicting defect-prone code reviews but to understand the impact our personal and participation metrics have on code review quality'
Adjusted R-squared of 0.123-0.173 across four models; the authors describe the predictive power as low
Backs overall reviewer experience only. A widely repeated module-specific version of this claim is stated in the paper as a hypothesis in its metrics table, not as a confirmed result
Oleksii Kononenko, Olga Baysal, Michael W. Godfrey2016 (ICSE 2016)data
A survey of 88 Mozilla core developers (22% response rate, 938 coded quotes, inter-coder agreement 94.2-97.2%) found review quality is primarily associated with thoroughness of feedback, the reviewer's familiarity with the code, and perceived code quality; 96% agreed reviewer experience influences review time.
Reviewers themselves identify expertise and code familiarity as the determinants of review quality, and name gaining familiarity with unfamiliar code as their single biggest challenge — corroborating the measured finding from the same group's quantitative study.
2 more excerpts
This is self-reported perception, not measured defect-detection data; it is the qualitative companion to the ICSME 2015 study and should be paired with it rather than substituted for it
Scoped to one large open-source community: 'While our findings might not generalize outside of Mozilla...'
Pereira, Sinha, Ghosh, Dutta (Nutanix, Inc.)10 Mar 2026data
code review agents can exhibit a low signal-to-noise ratio when designed to identify all hidden issues, obscuring true progress and developer productivity
"Find everything" review agents drown the signal, so resolution/merge rate is the wrong yardstick and signal-to-noise proxies developer trust.
Sebastian Baltes, Marc Cheong, Christoph Treude09 Jun 2026data
The development time has been shortened but the team now needs to spend more time to review. Doesn't look like any benefit.
Individual AI speedups externalize review burden onto the team, making review a shared, exhaustible resource rather than a free step.
3 more excerpts
"30 PRs per day across 6 reviewers" — [R07] reviewer-burden example
"reviewer-burden" ranked among top 3 most frequent codes (226 instances) — coding frequency
"Individual developers and organizations benefit from AI-generated content, but the cumulative effect degrades the shared resources that collaborative development depends on."
Shyam Agarwal, Courtney Miller, Christian Kästner, Bogdan Vasilescu (Carnegie Mellon University)July 8, 2026datapartial
A causal model of 26 constructs and 67 relationships (64 directed, 3 contested) built from 38,709 grey-literature documents, with a stratified sample of 3,100 coded via an LLM-assisted pipeline. Within the authors' own data, 40.1% of agent-authored PRs were examined only by the developer who invoked the agent, versus 21.5% of human-authored PRs.
Code review is the control point through which a coding agent's effect on software is decided, and the sign of that effect is set by the team — via reviewer expertise and how it structures review — rather than by the tool. The paper names three moderators: reviewer expertise and disposition, automated-reviewer capability, and process adaptation.
3 more excerpts
The authors demonstrate directional instability inside their own dataset: whether the developer who invoked an agent counts as an independent reviewer or as the author reviewing their own work decides whether independent review of agent PRs sits above or below the human rate
The reversal is scoped to the independent-review construct only; merge speed and discussion volume are presented as stable, non-reversing findings
Unrefereed preprint that self-describes as 'a proposed explanatory theory, not a validated one'; several figures cited elsewhere for this paper (4.6%/18.7%, 20.8%/29.3%, the Wilcoxon statistics) could not be located in the text across two independent verification passes and are not used
Across 278,790 code review conversations in 300 open-source GitHub projects, human reviewers' suggestions were adopted at 56.5% versus 16.6% for AI agents — a 39.9 percentage-point gap. Over half of unadopted AI suggestions (28.7% incorrect plus 24.0% superseded by an alternative fix) were not usable as written.
Automated-reviewer capability is a measurable quantity with a wide gap from human review, and adopting AI suggestions carries a quality cost: when adopted, agent suggestions produce significantly larger increases in code complexity and code size than human ones.
3 more excerpts
Over 95% of AI agent comments fall into Code Improvement and Defect Detection, while humans also supply Understanding, Testing, and Knowledge Transfer feedback
Results are pooled averages across 300 projects with no breakdown by project, team, or reviewer characteristics, and the authors note findings may not generalize to proprietary enterprise systems
Preprint; the AI-agent population mixes coding agents and purpose-built review bots (GitHub Copilot, CodeRabbit, Devin, Claude Code, Gemini Code Assist)
A survey of 2,989 developer responses plus 11 in-depth interviews at a single enterprise surfaced six productivity factors, two of them long-term — technical expertise and ownership of work — that per-sprint output metrics do not capture. Peer review is one of the six named factors.
What AI assistance puts at risk is a team-level capital stock (expertise, ownership, comprehension) rather than per-PR throughput, which is why merge-speed dashboards can move in a healthy direction while the thing that determines review quality erodes underneath them.
3 more excerpts
Interviewees emphasized 'the importance of growth as a developer over time, rather than just an individual's output per sprint'
Within the peer-review factor, a senior manager reports junior developers optimizing the wrong things in ways that cost others time to fix — the same tooling producing different outcomes by seniority inside one company
Single-company sample (BNY Mellon); code review is one of six factors rather than the study's focus
Argues that the purpose of review changes with a team's position along three variables — blast radius, how long the code lives, and how many people need to understand it — and that 'the only answer that survives contact with a real codebase is that it depends entirely on who you are.'
The most authoritative named-author treatment reaches contingency rather than a universal verdict, but segments on situational rather than team-capability variables and stops short of giving a reader any way to determine which regime their own team occupies.
4 more excerpts
Tier by risk, not by author. A config change earns a linter and a glance. A payments path earns the full stack
Carries the strongest counterargument to team-level explanations: 'teams with mature, disciplined engineering practices were hit just as hard as everyone else', citing Faros AI
Offers its own vendor-bias caveat: 'CodeRabbit and Faros both sell into this market, so their framing is not disinterested'
'We made writing cheap, and understanding stayed exactly as expensive as it has always been'
Andrea Griffiths (The GitHub Blog)May 7, 2026practitioner
GitHub's own platform telemetry: Copilot code review has processed over 60 million reviews, growing 10x in less than a year, and more than one in five code reviews on GitHub now involve an agent.
Agent involvement in review is now the normal case rather than an edge case, and the highest-authority platform guidance responds with standardized review practices applied uniformly, with no conditioning on team maturity, expertise distribution, size, or review culture.
2 more excerpts
Frames the core problem as invisible cost: 'The surface looks clean. The debt is quiet.'
Cites the January 2026 study 'More Code, Less Reuse' for the finding that agent-generated code adds redundancy and technical debt per change while reviewers feel better about approving it
The reviewer role is being automated. The review, understood as judgment about whether the software is correct for its purpose, is relocating to where the agent cannot follow.
Agents can take over diff inspection, but human judgment doesn't disappear — it relocates to intent specification up front and accountability at merge.
3 more excerpts
"An agent-assisted developer produces more pull requests per day than human review capacity can absorb." — Monperrus paper discussion
"Automate the checkpoint and the judgment does not evaporate. It relocates to intent specification on the way in and accountability on the way out"
"The human does not leave the loop. The human moves from the end of it to the start."
More code is entering the pipeline, but less of it is reaching production successfully. The bottleneck has moved from writing code to deciding whether code is safe to merge.
Third-party delivery data shows generation is not the wall — validation is, with feature throughput rising while main-branch throughput and success rates fall.
3 more excerpts
"feature branch throughput up 59% year over year, while main branch throughput for the median team actually fell" — CircleCI 2026 State of Software Delivery report
"main-branch throughput fell nearly 7%, and main-branch success rates dropped to 70.8%" — CircleCI 2026
"agentic AI PRs have a pickup time 5.3x longer than unassisted PRs. AI-assisted PRs wait 2.47x longer" — LinearB 2026 Software Engineering Benchmarks Report
The bottleneck moves from generation to review queues, CI capacity, flaky environments, branch policy, cost ceilings, and the human attention needed to decide what should actually merge.
As agents get capable, the constraint shifts off code generation and onto the whole delivery surface — review bandwidth, CI, and human merge decisions.
3 more excerpts
"The model matters, but the delivery surface matters just as much."
"A team that cannot write crisp tasks will struggle to evaluate agents honestly."
"Reviewers do not need another wall of generated explanation. They need the shortest path to deciding whether the change should merge."
Dex Horthy (HumanLayer)Jul 23, 2026 (undated in body; dated by commit history)practitioner
'no amount of harness engineering or loopsmaxxing can solve what is fundamentally a model-training issue.'
The strongest counterargument to a harness-centric thesis, and the reason this post bounds its claim to blast radius rather than quality.
4 more excerpts
On the limit of fast deterministic gates: 'Running the tests gets you a clean pass or fail in ~seconds... But the cost function of bad architecture is measured in weeks, months, maybe even years'
'if you build a harness but you don't own the weights and can't RL the model inside it, you'll always be at a disadvantage to a team that owns both'
Cites Faros AI: 31.3% of PRs skip review entirely, +242.7% incidents per PR under high AI adoption
The words determinism, guardrails, permissions, sandboxing, and policy enforcement appear nowhere in the document - it argues about design quality, not authorization
Based on roughly 5,000 survey responses and more than 100 hours of qualitative interviews, DORA frames AI's primary role as an amplifier that magnifies an organization's existing strengths and weaknesses, and locates the greatest returns in the underlying organizational system rather than the tools.
An independent method — survey and interview rather than telemetry — reaches the organizational-contingency conclusion, and stands in direct, named contradiction to Faros's telemetry finding that engineering maturity offered no protection.
3 more excerpts
Google Cloud-sponsored; headline figures are self-reported perception rather than measured delivery outcomes, and the amplifier thesis is close to unfalsifiable as stated
The full report is gated behind a lead-generation page and could not be read directly; the amplifier framing was verified on the public landing pages
The 2025 report does not address code review; a widely repeated 2024-to-2025 reversal on delivery throughput could not be verified and is not used
Synthesizing an interview with Morgan Stanley engineers, argues review shifted rightward onto scarce senior engineers, with Khalid Elsawaf noting a simple prompt could produce a thousand-file, 10,000-line pull request: 'No human here is going to review that.'
The canonical statement of the 'AI broke code review' position, asserted as a general industry pattern with no conditioning on team properties and no admission that the effect might run the other way for some teams.
2 more excerpts
Overt vendor content: the piece closes with a section on Moderne's own product and quotes Moderne's CEO; Morgan Stanley is a named Moderne customer
Anecdote-driven with no quantified metrics, surveys, or study citations
Shan Appajodu, Ravi Boyapati (Salesforce Engineering)January 29 (year not shown on page)practitioner
Internal signals reported without population, period, or sample size: code volume up approximately 30%, pull requests regularly exceeding 20 files and 1,000 lines, review latency rising quarter over quarter, and review time for the largest pull requests beginning to plateau or decline.
A worked example of a review metric reported rather than interrogated: the team asserts that plateauing review time on large PRs indicates reviewers disengaging, considers no alternative reading of the same signal, and never defines who counts as a reviewer.
2 more excerpts
Explicitly scoped to one organization rather than generalized
No publication year appears anywhere on the page, and no external research is cited
Argues the bottleneck has moved from writing code to deciding what should merge, and that dropping AI tooling into an unchanged lifecycle treats it as a productivity patch that will not last.
The freshest executive-facing coverage occupies the prescription slot rather than the diagnosis slot: it recommends a target state without offering a way to determine what is currently true of the reader's own team on expertise, reviewer capability, or policy adaptation.
2 more excerpts
It does carry one narrow genuine diagnostic — expand AI review scope only when suggestion acceptance rate is climbing and post-merge defect rate is flat or falling, treating one signal moving without the other as a red flag
Its statistics (CloudBees 2026 State of Code Abundance, GitHub Octoverse 2025, Stack Overflow 2025 Developer Survey) are cited second-hand and were not independently verified
When agents help vs. hurt; single vs. multi-agent; build vs. buy.
ConvAI InnovationsSeptember 2026datapartial
Jev figures are third-party published, never measured here (no TypeSafe API access); sample sizes and prompts differ. Before temperature scaling, the base checkpoint has higher raw ECE (0.213 vs 0.144).
The only structured non-vendor comparison of Jev is competitor-published and self-limiting, and its calibration advantage exists only after a temperature-fitting step that Jev did not receive.
4 more excerpts
Reported splits: Laya leads on typed-decisions (0.766 vs 0.727) and AG News (0.950 vs 0.910), while Jev leads decisively on Banking77 (0.870 vs 0.425), a high-option-count task
Latency p50 is reported as 32.8ms for Laya against 236-276ms for Jev
On DAIR Emotion the card reports Jev assigned zero probability to the true label on 16% of examples
An independent tester re-running on benchmarks with published Jev numbers found the two model cards' headline accuracy figures came from different benchmarks and were not comparable
David Klotz (IAAI, Media University Stuttgart)April 29, 2026 (arXiv:2604.26482v1)data
Mission-critical systems of record: Retain Buy as the primary option. Consider Make selectively for peripheral modules, extensions, or integration layers where the core system's integrity is not at risk.
Agentic AI shifts make-vs-buy by application type: commodity and differentiating apps move toward build, while regulated and mission-critical systems stay buy.
3 more excerpts
Commodity utilities: Default to Make. Evaluate Buy only where ecosystem integrations provide strong network value or where the firm's AI capability is below the viability threshold.
Where software development once required large teams working over months, small teams augmented by AI agents can now deliver functional applications in days or weeks.
AI-era Make demands skills in prompt engineering, agent orchestration, AI output validation, and governance of AI-generated artifacts.
David Wood (O'Reilly)August 2009 (book publication)data
Fully 60% of the life cycle costs of software systems come from maintenance, with a relatively measly 40% coming from development.
Maintenance dominates software lifecycle cost, and most of that maintenance is new enhancement work rather than bug-fixing.
1 more excerpt
During maintenance, 60% of the costs on average relate to user-generated enhancements (changing requirements), 23% to migration activities, and 17% to bug fixes.
DORA (Google Cloud)2024 (page last updated April 13, 2026)datapartial
AI adoption significantly increases individual productivity, flow, and job satisfaction. However, it also negatively impacts software delivery stability and throughput
AI helps the individual developer but hurts system-level delivery stability and throughput.
1 more excerpt
Unstable organizational priorities cause meaningful decreases in productivity and substantial increases in burnout.
GitClearJanuary 2026 (research notation on page)data
the percentage of changed code lines (associated with refactoring) sunk from 25% of changed lines in 2021, to less than 10% in 2024, while lines classified as 'copy/pasted' (cloned) rose from 8.3% to 12.3%
AI-assisted development correlates with more code duplication and less refactoring, increasing long-term maintenance burden on code you own.
3 more excerpts
211 million changed lines from repos owned by Google, Microsoft, Meta, and enterprise C-Corps
4x more code cloning
'copy/paste' exceeds 'moved' code for first time in history
Across 1,902 instrumented runs plus a 244-run sealed-environment replication, naming one agent as coordinator in its prompt creates no communication hub and no reliable improvement in success; shared files cut output tokens by about 42% at eight agents on message-heavy work.
The cheapest fix for multi-agent coordination failures, a role label in one agent's prompt, does not work; structural changes (shared artifacts) do.
2 more excerpts
The null result on coordinator-naming reproduced in the sealed 244-run replication, not just the primary 1,902-run set
Scope is a coding-agent benchmark with a fixed test suite and specific team configurations; the authors do not claim broader generality
Hoagy Cunningham, Alwin Peng, Jerry Wei, Euan Ong, Fabien Roger, Linda Petrini, Misha Wagner, Vladimir Mikulik, Mrinank Sharma2025data
Using Claude 3.5 Haiku as a safety filter for Claude 3.5 Sonnet increases inference costs by approximately 25%.
The tier a decision model displaces already has a published price: if a small LLM is doing your filtering today, that is roughly a quarter of your inference bill, which is the baseline a vendor's speed multiple is quietly measured against.
4 more excerpts
EMA linear probes outperform, at negligible cost, a dedicated classifier with 2% of the policy model's parameters
Two-stage pipelines reduce cost by over 10x without significantly reducing overall system performance
The page states performance at low false-positive rate is the production-relevant metric, because that determines suitability for deployment
The page carries no publication date; only the 2025 path segment in the URL
Hoagy Cunningham, Jerry Wei, Zihan Wang, Andrew Persic, Alwin Peng, Jordan Abderrachid, Raj Agarwal, Bobby Chen, Austin Cohen, Andy Dau, Alek Dimitriev, Rob Gilson, Logan Howard, Yijin Hua, Jared Kaplan, Jan Leike, Mu Lin, Christopher Liu, Vladimir Mikulik, Rohit Mittapalli, Clare O'Hara, Jin Pan, Nikhil Saxena, Alex Silverstein, Yue Song, Xunjie Yu, Giulio Zhou, Ethan Perez, Mrinank SharmaJanuary 8, 2026data
Since exchanges flagged by the first stage are escalated rather than refused, the first-stage classifier can flag a higher proportion of production traffic without incurring an excessive refusal rate. The first-layer probe escalated approximately 5.5% of traffic to the second-stage classifier, for approximately a 40x reduction compared to the single exchange classifier.
Escalating rather than refusing changes which false-positive rate you can tolerate, which is why the band you choose matters more than the accuracy number the vendor quotes.
4 more excerpts
The refusal flag rate was 0.05%, down from the 0.38% reported for the prior system
1,736 cumulative hours of red-teaming effort across approximately 198K attempts
Both systems are trained on synthetic data using a constitution related to CBRN weapons, so the domain is narrower than general-purpose triage
An earlier research pass recorded the escalation rate as roughly 10 percent; the paper states approximately 5.5 percent
Kanghee Park, Jiayu Wang, Taylor Berg-Kirkpatrick, Nadia Polikarpova, Loris D'Antoni2024data
Constrained decoding techniques can distort the LLM's distribution, leading to outputs that are grammatical but appear with likelihoods that are not proportional to the ones given by the LLM, and so ultimately are low-quality.
Guaranteeing that an output conforms to a schema is not the same as guaranteeing the output is right; constraining the answer space can actively degrade which valid answer you get.
1 more excerpt
The authors propose grammar-aligned decoding (ASAp) to preserve grammaticality while matching the model's conditional distribution, so the distortion is a known and addressable artifact rather than an inherent cost
Kilian Hendrickx, Lorenzo Perini, Dries Van der Plas, Wannes Meert, Jesse DavisJuly 23, 2021data
Machine learning models always make a prediction, even when it is likely to be inaccurate. This machine learning sub-field was already studied in 1970 by Chow and Hellman.
Banding a model's probability with an operator-chosen threshold is the textbook reject option with a fifty-year literature, not a new capability introduced by a decision model.
3 more excerpts
The canonical formalism is a single threshold with two outcomes, so a three-band act/review/escalate structure is an extension of it rather than the base model
The standard cost ordering Cc < Cr < Ce carries a second constraint frequently dropped in summary: Cr must be no greater than 1/K for K classes
Ambiguity rejection and novelty rejection need different confidence functions: a class-posterior score signals ambiguity, while novelty needs a density or distance measure against the training distribution
Maksym Nechepurenko, Pavel ShuvalovMay 5, 2026data
Naming and isolating the coordination layer moves four classes of properties from implicit-in-code to explicit-in-spec (failure-mode signatures, cross-system comparability, agent heterogeneity), and the specification can be implemented atop AutoGen, CrewAI, LangGraph, AWS Strands, or Microsoft Foundry without modification.
Coordination logic should be treated as its own configurable, separable architectural layer rather than embedded in agent prompts, because separation is what makes the layer analyzable at all.
3 more excerpts
The paper's own experiment (five coordination configurations, 100 post-cutoff Polymarket markets, Murphy decomposition) tests forecasting calibration signatures, not production failure rates
The commonly-cited 41-87% production-failure-rate figure appears in this paper's introduction as a citation to Cemri et al. 2025, not as this paper's own finding
Pairwise comparisons between coordination configurations do not survive Bonferroni correction at n=100; the authors call this a methodology-validating first instantiation, not a general cross-model claim
Mrinank Sharma, Meg Tong, Jesse Mu, Jerry Wei, Jorrit Kruthoff, Scott Goodfriend, Euan Ong, Alwin Peng, Raj Agarwal, Cem Anil, Amanda Askell, Nathan Bailey, Joe Benton, Emma Bluemke, Samuel R. Bowman, Eric Christiansen, Hoagy Cunningham, Andy Dau, Anjali Gopal, Rob Gilson, Logan Graham, Logan Howard, Nimit Kalra, Taesung Lee, Kevin Lin, Peter Lofgren, Francesco Mosconi, Clare O'Hara, Catherine Olsson, Linda Petrini, Samir Rajani, Nikhil Saxena, Alex Silverstein, Tanya Singh, Theodore Sumers, Leonard Tang, Kevin K. Troy, Constantin Weisser, Ruiqi Zhong, Giulio Zhou, Jan Leike, Jared Kaplan, Ethan PerezJanuary 31, 2025data
These classifiers also maintain deployment viability, with an absolute 0.38% increase in production-traffic refusals and a 23.7% inference overhead.
A synthetic-data-trained classifier was already on the production hot path eighteen months before Jev launched, and its owners published the tax rather than only the accuracy.
4 more excerpts
The classifiers were trained on synthetic data generated by prompting LLMs with natural language rules, i.e. a constitution, which is the same training lineage a synthetic-only decision model sits in
Over 3,000 estimated hours of red teaming across 405 HackerOne participants, with bounties up to $15K per jailbreak report
The prototype's prioritization of robustness led to impractically high refusal rates, which is the tradeoff the next generation was built to fix
This sits at the input/output safety-filter layer rather than a general-purpose routing tier, so the analogy is about tax disclosure precedent rather than an identical architectural slot
if AI adoption increases by 25%, estimated throughput delivery is expected to decrease by 1.5%
Individual AI productivity gains do not translate into system-level delivery throughput or stability, because code generation was never the bottleneck.
3 more excerpts
estimated delivery stability is expected to decrease by 7.2%
75.9% of respondents (of roughly 3,000 people surveyed) are relying on AI for at least part of their job responsibilities
if AI adoption increases by 25%, time spent doing valuable work is estimated to decrease 2.6%
Sahana Chennabasappa, Cyrus Nikolaidis, Daniel Song, David Molnar, et al.May 6, 2025data
Approximately 90% of inputs are fully resolved by the first layer, maintaining a typical end-to-end latency of under 70 milliseconds; for the remaining 10% requiring deeper inspection, end-to-end latency can exceed 300 milliseconds. PromptGuard 2 runs at 19.3ms (22M) and 92.4ms (86M).
A production agent stack already layers a deterministic tier, a small-classifier tier, and an LLM auditor, each with a measured latency, which is the real comparison set for a new decision model rather than a frontier LLM.
4 more excerpts
The first tier's own scan time is approximately 60 milliseconds, which is a separate measurement from the under-70ms end-to-end figure for the 90 percent it resolves
Named tiers: CodeShield for static analysis over 50+ CWEs in seven languages, PromptGuard 2 as the DeBERTa-based classifier, and AlignmentCheck as a chain-of-thought LLM auditor
Threshold selection is reported as a fixed minimal utility reduction of 3% in the AgentDojo evaluation, which is a separate evaluation from the recall-at-1%-FPR classifier table
Role-scoped scanning: PromptGuard analyzes only user and tool messages while AlignmentCheck evaluates assistant messages
Sheryl Estrada (Fortune)August 18, 2025, 6:54 AM ETdata
Purchasing AI tools from specialized vendors and building partnerships succeed about 67% of the time, while internal builds succeed only one-third as often.
Most enterprise GenAI builds fail; buying and partnering succeeds roughly three times more often than building internally.
3 more excerpts
95% failure rate for enterprise AI solutions
about 5% of AI pilot programs achieve rapid revenue acceleration
150 interviews with leaders, a survey of 350 employees, and an analysis of 300 public AI deployments
Classical calibration metrics include expected calibration error, reliability diagrams, and proper scoring rules. These both provide static calibration assessment but do not address sequential monitoring with false alarm control.
Measuring calibration once is not monitoring it, and checking it repeatedly with a naive fixed-sample test manufactures false alarms, so threshold drift needs a sequential instrument rather than a recurring spot check.
1 more excerpt
A practitioner checking calibration daily at p<0.05 will almost certainly observe spurious alarms over a year even if calibration remains stable, because classical hypothesis tests assume a fixed sample size determined before seeing data
Such confidence-based deferral often works remarkably well in practice, but post-hoc deferral significantly improves on it where downstream models are specialists, samples carry label noise, or there is distribution shift between the train and test set.
A confidence threshold tuned in evaluation is not guaranteed to survive production traffic drift, which is the named condition under which routing on a cheap model's confidence stops being sufficient.
2 more excerpts
The paper's position is that confidence deferral is usually fine with three named exceptions, not that it is generally broken
The Bayes-optimal rule is a difference over both models' correctness probabilities against the deferral cost, so it requires the downstream model's expected correctness, which the first model's raw confidence score alone cannot supply
Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D Sculley, Sebastian Nowozin, Joshua V. Dillon, Balaji Lakshminarayanan, Jasper Snoek2019datapartial
We find that traditional post-hoc calibration does indeed fall short, as do several other previous methods.
Calibration established on one distribution does not automatically hold when the distribution moves, which is why a vendor's calibration figure has to be re-derived on your own traffic rather than adopted.
3 more excerpts
Population is 2019 image, text and tabular classifiers under synthetic corruptions, with no LLMs and no synthetic-training-data regime, so applying it to a decision model is an inference and is labelled as such in the post
The paper's headline positive result is that ensembles stay well calibrated under the same shift, so calibration is joint over model and distribution
A quote carried by an earlier automated research pass could not be located in the paper and was discarded; only the abstract sentence above is used
Using our method an unprecedented 2% error in top-5 ImageNet classification can be guaranteed with probability 99.9%, and almost 60% test coverage.
Buying a guaranteed risk level costs coverage, and the price is large and measurable: roughly 40 percent of inputs go unanswered to hold a 2 percent error bound.
2 more excerpts
Coverage is formally the probability mass of the non-rejected region, and selective risk is defined only over the accepted region normalized by coverage
CIFAR-10 reaches 1% error at 78.56% coverage and CIFAR-100 18.85% error at 67% coverage, both at delta=0.001
Yuanbo Zhou, Changjia Zhu, Junyu Wang, Xu He, Yan Zhai, Kun Sun, Mingkui Wei, Junjie XiongMay 22, 2026data
We identify a critical blind spot arising from the mismatch between the limited inspection windows of guardrail models and the substantially larger context inference windows of downstream LLMs.
A decision model's request budget is an attack surface rather than a quota: when the inspector's window is far smaller than the actor's, the gap itself is the vulnerability.
3 more excerpts
Guardrails with an effective 512-token window sit in front of models accepting up to 400k tokens, with interleave-layout bypass rates consistently above 99.5% against Prompt Guard 86M
The population is lightweight prompt-injection and toxicity classifiers, not calibrated multi-class decision models, so this is a reason to interrogate a decision model's budget rather than a measurement of one
Gemini 3 Pro recognized malicious intent in all 200 tested overflow cases, raising the question of whether the actor is a better last-line detector than an undersized upstream guard
Of the 253 actions, 93 (36.8%) were performed via Edit or Write tool calls, which are routed through Tier 2 and are not evaluated by the classifier.
A decision gate's coverage is a distinct failure channel from its accuracy: actions that never reach the model cannot be scored by it at all, and the gate's boundary has to match the actor's real action space rather than its most common channel.
3 more excerpts
The authors explicitly state the 81.0% stress-test FNR is not directly comparable to Anthropic's reported 17% on organic production traffic, calling the delta workload sensitivity rather than a flaw in the vendor's measurement
Coverage is the minority cause, not the dominant one: 51 of 115 false negatives came from the coverage gap and 64 from the classifier misjudging actions it did evaluate
Scope is a single model (Sonnet 4.6), one threat category, four synthetic DevOps task families with shimmed CLIs; the paper never mentions hallucination or schema conformance
A2A's opacity principle lets agents collaborate without sharing internal memory, proprietary logic, or specific tool implementations.
A shipping, cross-vendor-governed protocol already architects coordination as a separable layer via explicit design intent.
1 more excerpt
Cited as design intent only, not adoption evidence: no adopter case studies or production-deployment list found on the page, and the roadmap still lists authorization schemes, dynamic capability negotiation, and streaming reliability as unresolved
'Planning, implementation, and testing of the same feature share too much context' to split across agents, and 'Components requiring constant back-and-forth belong in the same agent.'
There are principled places not to cut the graph - shared context and high synchronization needs are the signals to keep work in one node.
1 more excerpt
Used in the post as the 'where not to cut' check in the pricing list, a counterweight to over-decomposition
Because flagged exchanges are escalated to the more powerful model, rather than refused, the first-stage classifier can afford a higher false-positive rate and not frustrate the user with refusals. Runs at just ~1% additional compute cost.
The cost of the cheap decision tier is not fixed: the same vendor repriced it from 23.7 percent overhead to roughly 1 percent in a year, so a vendor's cost comparison is a snapshot of a moving baseline.
3 more excerpts
Refusal rate of 0.05% on harmless queries, an 87% drop from the original classifiers system
The classifiers were trained on synthetic data generated from a constitution of natural language rules
The scoping qualifier about Claude Sonnet 4.5 traffic over one month was not isolated verbatim during verification and should be re-checked before being quoted
Bryan Ross (GitLab, Field CTO)March 24, 2026practitionerpartial
For a team of roughly 200 developers, an internal build typically costs around $1.4M in year one, requires 2–3 dedicated FTEs to maintain, and takes 12–18 months to reach a first real use case.
Building an internal agentic AI platform in regulated industries is a multi-year, multi-FTE commitment with governance surface most organizations underestimate.
2 more excerpts
Every engineer building the platform is an engineer _not_ modernizing a legacy pipeline, remediating security debt, or accelerating a critical delivery program.
Building an internal agentic AI platform in banking or insurance is a multi-year platform engineering commitment with regulatory surface area most organizations underestimate
Jev is TypeSafe's structured evaluation model. It evaluates one state against typed Noul, Choice, and Score questions and returns calibrated answers with probabilities and confidence. Context window listed as 32,000 tokens.
A third host independently describes the model's contract and publishes a 32,000-token context window, corroborating the cross-host discrepancy on the request budget.
1 more excerpt
Cloudflare does not publish per-token rates on the model page, deferring to the dashboard
CrewAI assigns coordination strategy as a named process type at crew creation, with the hierarchical process requiring its own manager_llm or manager_agent.
Coordination strategy is a swappable, crew-level configuration decoupled from individual agent and task definitions in a shipping framework.
The response reports the versioned ID that answered, so log it, and pin that ID once you've tuned thresholds against it.
Threshold drift has a version axis: a floating alias moves the model your thresholds were tuned against, and nothing in your logs records that it happened unless you pin and log the resolved ID.
4 more excerpts
He reads the vendor's numbers carefully, noting the docs use 0.5 as a review floor and 0.9 before a destructive action as examples rather than defaults
He documents hands-on accuracy limits: unreliable at counting, comparing numbers, and date or window comparisons, with indirection and large irrelevant state both degrading accuracy
He notes scores are not a measurement and are usable to threshold or rank rather than to interpolate magnitude
The origin returns HTTP 403 to direct fetches and web.archive.org was blocked in this environment; the source was recovered via a reader proxy
A storage-agnostic handoff contract (objective and ownership, a five-state machine, observations vs. mutations, verification evidence, approval gates) needs no new service, a JSON file, a database row, or a session handoff can carry it.
A concrete, lightweight artifact is what a separated coordination layer looks like in practice, not a new orchestration platform.
1 more excerpt
Single-author field pattern, not peer-reviewed or independently corroborated; paired with Cemri et al.'s Task Verification failure category as the academic anchor for the verification half of the claim
Jev averages 67.8% agreement with the reference answers, where the reference answers are themselves average predictions of GPT-6 Astra and Anthropic's Fable.
The headline accuracy figure measures agreement with a synthetic consensus of two other vendors' models rather than with ground truth, which is a weaker claim than accuracy parity.
3 more excerpts
The piece states all numbers come from TypeSafe's own evaluation tables with no large-scale independent reproduction, and should be treated as vendor-reported
The workflows were written by TypeSafe's own model capabilities team, which the author flags as a possible source of bias
Comparison figures cited: GPT-5.6 Terra at 67.9% accuracy and 10.1s, Claude Opus 5 at 73.1% and 37.8s, with Jev reporting 0% structured output error
Turn order in AutoGen's group-chat pattern is maintained by a separate Group Chat Manager agent, implemented as its own class, configurable as round-robin or LLM-selected.
Shipping open-source frameworks already extract coordination logic (turn-taking, speaker selection) into a separate, configurable component distinct from individual agent definitions.
Treat moderation scores as signals for your application's policy, not as an automatic blocking decision. We plan to continuously upgrade the moderation endpoint's underlying model. Therefore, custom policies that rely on category_scores may need recalibration over time.
A vendor shipping this exact product shape, a per-category confidence score plus a typed flag, has been telling customers for years that the threshold is theirs and will need recalibration when the model moves.
2 more excerpts
The endpoint is free to use, so the precedent is not confounded by pricing pressure
The docs never use the words alias or floating for omni-moderation-latest; that characterization is the post's analysis, supported in substance by the recalibration warning
OpenAI uses small, high-recall classifiers to determine if content is domain-relevant to priority risks before evaluating that content with gpt-oss-safeguard. Traditional classifiers have lower latency and cost less to sample from.
The cheap classifier tier and the expensive reasoning tier are complements rather than substitutes, stated by a vendor that ships both.
4 more excerpts
Traditional classifiers trained on thousands of examples will likely perform better on a task than the reasoning model
Rules alone handle deterministic cases well, such as keyword matches and metadata thresholds, but struggle with satire, coded language, or nuanced policy boundaries
The guide advises pre-filtering content sent to the reasoning model because it is more time and compute intensive than other classifiers
No clean general statement exists on the page instructing escalation of ambiguity to a human; only policy-template example text
typesafe/jev-1.13 lists input at $0.042/M and output at $0/M with 32K context; the jev-latest entry states it always redirects to the latest model in the Jev family.
An independent host confirms both the pricing and the floating-alias behaviour, and publishes a context figure that does not match TypeSafe's own documentation.
1 more excerpt
TypeSafe's own Models page states 64k tokens per request with 32k for state plus the longest question, while OpenRouter and Cloudflare both publish 32K, so the number a reader designs against depends on which host they read
Pat Brans (CIO.com)December 11, 2025practitionerpartial
With such a layer in place, the build-versus-buy question fragments, and CIOs might buy a vendor's persona agent, build a specialized risk-management agent, purchase the foundation model, and orchestrate everything through a platform they control.
The industry consensus has shifted to hybrid: assemble build and buy across the AI stack under an orchestration layer you control.
2 more excerpts
Six months ago many were experimenting, but now they're scaling.
including cases where a senior executive's data surfaced in a junior employee's query.
To me, this seems like a semantic dodge, since Jev can absolutely still pick the wrong choice (e.g. calling the sky "red"). Still, all of this is also true about regular LLMs with structured outputs, and it doesn't make Jev any more reliable in practice.
The zero-hallucination claim is about schema conformance, a property any LLM with structured output already has, and it says nothing about whether the decision is correct.
4 more excerpts
He reports no evidence the returned probabilities are anything other than regular logit probabilities
His own test on Qwen2.5-1.5B-Instruct with prefix plus constrained single-token decoding produced a 2x-3x speedup, suggesting the speed is an inference-stack result rather than a new architecture
Not being able to use test-time compute is a real ceiling that likely caps this model class around the strength of non-reasoning LLMs
His position is that Jev is useful but not architecturally novel, not that it is worthless
For serious work, a specific hand-built classifier will always be cheaper and faster than Jev. Generic classifiers have to encode knowledge of all kinds of irrelevant things in their weights, so they can address lots of different tasks. That makes them larger, slower, and more expensive to run.
A generic decision model may earn its slot only while you are still unsure the feature works, because its own logged input and output pairs become the training set for the bespoke classifier that replaces it.
3 more excerpts
The post is explicitly speculative, framed as what the author expects to become a common pattern, so a finite shelf life is a prediction rather than established economics
The bespoke path still requires ML expertise, merely deferred until after the feature is validated, so what is removed is the dataset obstacle and not the skill obstacle
The post says nothing about version pinning, rate limits, API shape or ownership economics
The factor that stands out most to me is that these developers were all working in repositories they have a deep understanding of already, presumably on non-trivial issues since any trivial issues are likely to have been resolved in the past.
AI's edge is smallest exactly where you own and deeply understand a mature codebase long-term.
3 more excerpts
56% had never used Cursor before the study
Developers accepted less than 44% of AI generations
A quarter of the participants saw increased performance, 3/4 saw reduced performance
I can only focus on reviewing and landing one significant change at a time, but I'm finding an increasing number of tasks that can still be fired off in parallel without adding too much cognitive overhead to my primary work.
Human review-and-land throughput — one significant change at a time — is the real ceiling on how far parallel agents scale.
1 more excerpt
Code that started from your own specification is a lot less effort to review.
At the end of the day, it delegates the hallucination problem a little bit to the user. The user has to say, okay, if this only comes back with 50% probability, maybe this is a coin toss, and I disregard it. But if it's 95%, sure, then I can do something with it.
Buying a calibrated decision model transfers the judgment, rather than removing it: the vendor guarantees the shape of the answer and the operator inherits the threshold that decides what to do with it.
2 more excerpts
The quote is Armin Ronacher, CTO of Earendil; no first-party post on Jev exists on his own site, so this article is the primary of record
Diogo Almeida on the training approach: an early bet that we will be making all of our data, described as one of the best bets he has made
Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct.
The vendor's own documentation limits the claim its launch post and the downstream coverage make: a calibrated probability is an aggregate property, so a single Jev answer can still be wrong.
3 more excerpts
Defines the class: System One models are built to make fast, structured decisions software can use directly, and do not write replies, produce code, or explain their reasoning
Confidence is returned so the caller can decide when to act and when to escalate to a person or a reasoning model
The page does not define RLCD and does not state cost-scaled confidence bands, despite both being attributed to TypeSafe across the coverage
An alias moves when a new release ships, so the answers behind it can change without a change on your side. jev-latest resolves to jev-1.13.0; 64k tokens per request with 32k for state plus the longest question; 250,000 tokens per second and 1,200 requests per minute; input $0.042/M, output free.
Adopting a hosted decision model behind a floating alias means the thing your thresholds were tuned against can change with no deploy on your side.
3 more excerpts
Rate limits are adjusting dynamically and the published limits can change without notice
Text only at launch: string, JSON object, or array of text values, with no image, audio, or video input
TypeSafe publishes 64k per request while Cloudflare and OpenRouter both publish 32K context, so the hosts do not agree on the number a reader would design against
Test thresholds by plotting confidence against accuracy on your data.
The vendor itself instructs callers to derive thresholds locally rather than adopt a published number, which is the strongest available support for treating the threshold as operator-owned.
3 more excerpts
Guidance is two-outcome: make code take different actions for confident and unconfident answers, escalating uncertain cases to a person or a more expensive reasoning model
Code examples show a two-sided uncertainty window (0.4 < spam_risk < 0.6) and a single floor (confidence < 0.8 routes to human review)
The docs never state cost-of-being-wrong scaling; that framing is the post's extension and is labelled as such
POST https://api.typesafe.ai/v1/systemone takes state plus a map of named typed questions and returns answers keyed by question id; the response echoes the resolved version (jev-1.13.0) even when the request said jev-latest.
The integration is not a drop-in model swap: the request and response shape diverge from the OpenAI chat-completions convention, and the resolved version is visible in the payload, which is what makes pinning auditable.
2 more excerpts
Noul returns a 0-1 probability; Choice returns the selected option plus a probabilities map plus confidence; Score returns a weighted value plus a legend
The docs never themselves draw the contrast with OpenAI's shape; that comparison is the post's observation from the two schemas
The model never makes type errors. Claims 40x-200x faster, 70ms-500ms end-to-end, $0.042/MTok input with free output, and a homepage figure of 193.6x faster and 444.6x cheaper.
The launch page's absolute claim sits against the same vendor's concepts page conceding calibration does not guarantee an individual answer is correct, so the zero-hallucination framing is about schema conformance, not correctness.
2 more excerpts
Benchmark provenance is vendor-run against a vendor-selected comparison: numbers come from Workflow evals testing how they compare to the average of the smartest models, in this case Astra and Fable
The page does not state training exclusively on synthetic data and does not define RLCD, despite both being widely attributed to it
Two named principles: 'Share context, and share full agent traces, not just individual messages' and 'Actions carry implicit decisions, and conflicting decisions carry bad results.' Illustrated with a Flappy Bird clone split across two subagents whose outputs could not be reconciled.
A written handoff specification cannot capture the implicit decisions an agent makes mid-task, which is why parallel subagents without shared context produce conflicting assumptions that were never prescribed upfront.
4 more excerpts
'Actions carry implicit decisions, and conflicting decisions carry bad results.' On the parallel-subagent failure: 'The actions subagent 1 took and the actions subagent 2 took were based on conflicting assumptions not prescribed upfront.'
The essay's actual position is broader than an argument about edge specification: 'in 2025, running multiple agents in collaboration only results in fragile systems.' It is the steel-manned opposition to any multi-agent topology, not a supporting voice.
Its proposed alternative is a single-threaded linear agent.
Cognition's April 2026 follow-up states that multi-agent systems work best when writes stay single-threaded and additional agents contribute intelligence rather than actions