Field Notes from Running AI Agent Teams in Production
One studio, AI agent teams, real costs, honest failures. Most AI development content is vendor marketing or tutorials. This is what it actually looks like to ship products with agents. The agent workforce runs on Claude Code - Anthropic's CLI harness - with Claude as the foundation model.
New here?Start Here for curated reading paths, or readThe System for the operating model behind it.
Articles
Deep dives on a single topic.
Our Grades Fell While the Code Got Better
Five scored code reviews trended downward while each one triggered a real remediation wave. The score was measuring the reviewer, not the code.
Read article →Sixty Notes Nobody Could Read
An audit found 60 saved agent memories reachable from no index at all. The notes were fine. The delivery tier they were wired into was the defect.
Handing a Decision to a Human Needs a Real Acknowledgment
An escalation ending in "reply with this code" is empty unless the code names a stable item and the reply names a verified person.
The Watchdog That Announced a Restart It Never Performed
A supervision mechanism that logs its intent and then fails to act is worse than no supervision, because the log reads like recovery in progress.
Ship Log
What changed this week.