
I’m John Young.
I run engineering at Combine Capital and Infrared, and write research-backed essays on running AI coding agents in production — every claim traced to a primary source.
github.com/johnayoung linkedin.com/in/jyoung1985 john.anto.young@gmail.com rss
Ledger
Essays24
Pillars6
Evidence base verified2026-09-07
Recent Essays
- Diff Review Can’t See Your Agent Regress. A Standing Suite Can. Evals & Verification Per-diff review checks one PR, not whether your coding agent still works. How to run AI agent regression testing as a standing eval suite: triggers, tiers, owners.
- Your Agent’s “Done” Is a Self-Report. Make It a Predicate. Task Design Acceptance criteria for AI coding agents fail when done is prose the agent grades itself. Write end-state predicates a checker runs where the agent can’t write.
- Where a Decision Model Belongs: Placing Jev in Your Stack Architecture Decisions Jev’s launch numbers answer a question you don’t have. The tier it joins already ships at Anthropic, Meta and OpenAI, with the costs published.
- Coordination Is an Architecture Layer, Not a Prompt Instruction Architecture Decisions When a multi-agent pilot breaks, the postmortem blames the model. The failure data says coordination logic buried in each agent’s prompt is the real defect.
- AGENTS.md vs CLAUDE.md: Swap the Agent, Keep the Harness Agent Runtime Claude Code reads CLAUDE.md, not AGENTS.md. Sort your harness into what ports across vendors and what does not, so switching coding agents is a config flip.