Subscribe
Sign in
Home
Notes
Start here
Loop Lab
Podcast
Newsletter
Archive
About
Latest
Top
What we've learned (so far) making CircleCI work for agents
Five things we learned about context, tools, and evaluation when we started building for AI agents instead of humans
15 hrs ago
•
Confident Commit
2
Bake low Merge Efficiency Ratio (MER) into the agent loop
How engineering your agent loop for low MER (with inner-loop checks, a CI sneak peek, and a single push) cuts outer-loop CI credits by 5–6x without…
Sep 22
•
Ryan E. Hamilton
1
The only context LLMs can’t fake
Rob Zuber sits down with Dennis Pilarinos to talk about what agents don’t know and why it’s slowing your team down.
Sep 15
•
Confident Commit
39:51
We built a new API for AI agents. None of them used it (directly).
Agents make 97% of their calls against our new agent-friendly API. Not one agent found the API by reading the catalog, the llms.txt or the OpenAPI spec.
Sep 15
•
Dan Mullineux
Your CLI docs are too long. Claude stopped reading them.
The CircleCI CLI had good help text. Claude was only reading the first 40 lines of it.
Sep 15
•
Pete Steyert-Woods
I ran 10,000+ configs through the compiler in 4 Minutes
How a fleet of 100 Chunk sidecars turned a 10-hour sequential validation job into a 4-minute answer for a customer facing a breaking change.
Sep 11
•
Makoto Mizukami
1
1
Pencils down: The cost of a green PR
Giving a coding agent a pre-run inventory of exactly what CI will check eliminated fix loops, driving green commit rates to 100% and rework cost to $0…
Sep 11
•
Ryan E. Hamilton
2
1
August 2026
We cut 25% of tokens fixing CI. Here’s how
It used to be that an agent, when handed a failed CI run, spent many turns just figuring out what broke. Here’s how we changed that.
Aug 31
•
Joaquin Sandoval
1
1
I asked Grok 4.6 and Sonnet 5 to fix failed CI pipelines. CircleCI was the judge.
Grok 4.6 went 6 for 6 fixing broken CI pipelines at $0.22 a fix; Claude Sonnet 5 at max effort went 3 for 6 at $1.11 — and the expensive model's…
Aug 26
•
Ryan E. Hamilton
1
2
Three habits of teams shipping 9x more validated code
Yesterday, we published our Q2 Pulse report, a follow-up to our annual State of Software Delivery that digs into a question engineering leaders are…
Aug 18
•
Confident Commit
1
Are you paying too much to ship?
Earlier this year, the conversation in software was all about maximizing agent output. Tokenmaxxing. Jensen Huang telling engineers to spend $250K a…
Aug 18
•
Confident Commit
1
Does VCS choice determine your delivery destiny?
You can tell a lot about a developer from the tools they use.
Aug 18
•
Confident Commit
1
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts