Field Notes
A course-style series on running Claude Code and agentic pipelines in production, built entirely from incidents on this site's own infrastructure — what broke, why it was hard to see, and the actual fix.
For engineers already running Claude Code and agentic pipelines in production.
Cost & Infrastructure Economics
What it actually costs to run Claude Code day to day and on a schedule — and the specific changes that cut real spend without cutting quality.
- 1The Cheapest Way to Run Scheduled Claude Jobs Isn't the API→
- 2The Env Var That Secretly Tripled Our AI Coding Bill→
- 3When Prompt Caching Costs You More Than It Saves→
- 4Why Your Spend Limit Doesn't Survive a Fresh CI Checkout→
- 5How Five Free Workflows Added Up to a Real BillComing soon
- 6Stop Letting Your AI Agent Reroll From Scratch After Every RejectionComing soon
Silent Failures
The absence of an error is not the same claim as "it worked." Seven incidents where a pipeline looked healthy while doing nothing, or the wrong thing.
- 7The Page That Loaded Fine and Showed NothingComing soon
- 8The Eighteen Tests That Never Actually RanComing soon
- 9We Built a Drift Detector That Could Never Detect DriftComing soon
- 10Your Alerts Need Their Own AlarmComing soon
- 11The Outage Where Every Health Check Said Everything Was FineComing soon
- 12A Feature Being Off Looks Exactly Like a Feature Being BrokenComing soon
- 13One in Nine of Our Articles Was Cut Off Mid-SentenceComing soon
Credentials & Secrets
Agentic workflows change how credentials get created, checked, and leaked. Four incidents that weren't about a stolen key — they were about a check that lied.
- 14A Redaction Script Hid Every Secret Except OneComing soon
- 15Why Our Health Check Said a Working Key Was DeadComing soon
- 16The Error Message That Sent Us In the Wrong DirectionComing soon
- 17The Day Our Three Secret Stores DisagreedComing soon
Running Claude Unattended
Scheduling agentic work and coordinating more than one agent surfaces failure modes interactive sessions never hit.
- 18Our Local Cron Jobs Kept Missing Their RunsComing soon
- 19The Scheduled Jobs Nobody Told the Scheduler AboutComing soon
- 20A Green Checkmark Doesn't Mean the Work Got DoneComing soon
- 21What Happens When Two AI Agents Race to Log the Same EventComing soon
- 22The Auto-Merge Setting That Doesn't Actually Protect AnythingComing soon
Deploy, CI & Git
Shipping what an agent builds hits its own class of problems — most of them about ordering, not code.
- 23The Deploy Command That Broke Three Different Ways in One MonthComing soon
- 24A Squash Merge That Changed Nothing and Broke EverythingComing soon
- 25Our Security Scanner Got Cancelled by the Step Before ItComing soon
- 26The Webhook That Was Faster Than Its Own DatabaseComing soon
- 27One Leftover File Broke Every CI Run for a WeekComing soon
Running Claude Code in production is a program, not a prompt
If your team is scaling from individual Claude Code use to agentic pipelines running unattended, most of the failure modes above show up eventually. Happy to talk through what we've built.
Work with me →