<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Confident Commit]]></title><description><![CDATA[The conversations, data, and real-world experiments shaping how software gets shipped, from the people studying it, leading it, and living it.
]]></description><link>https://www.confidentcommit.com</link><image><url>https://substackcdn.com/image/fetch/$s_!nIQR!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcfc0382e-96b7-4183-86b5-c2d29838b081_1024x1024.png</url><title>Confident Commit</title><link>https://www.confidentcommit.com</link></image><generator>Substack</generator><lastBuildDate>Thu, 24 Sep 2026 03:33:05 GMT</lastBuildDate><atom:link href="https://www.confidentcommit.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[CircleCI]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[circleci@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[circleci@substack.com]]></itunes:email><itunes:name><![CDATA[Confident Commit]]></itunes:name></itunes:owner><itunes:author><![CDATA[Confident Commit]]></itunes:author><googleplay:owner><![CDATA[circleci@substack.com]]></googleplay:owner><googleplay:email><![CDATA[circleci@substack.com]]></googleplay:email><googleplay:author><![CDATA[Confident Commit]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[What we've learned (so far) making CircleCI work for agents]]></title><description><![CDATA[Five things we learned about context, tools, and evaluation when we started building for AI agents instead of humans]]></description><link>https://www.confidentcommit.com/p/designing-for-agent-experience-ax-lessons-from-circleci</link><guid isPermaLink="false">https://www.confidentcommit.com/p/designing-for-agent-experience-ax-lessons-from-circleci</guid><dc:creator><![CDATA[Confident Commit]]></dc:creator><pubDate>Wed, 23 Sep 2026 13:01:24 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a676ea47-6cf9-4030-b849-afaa989ed7fd_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Software companies have spent decades designing products around the needs of the people who use them. We started by focusing on </span><strong><span>user experience (UX)</span></strong><span>: making software understandable, intuitive, and effective for end users. Then, as software became something developers increasingly had to integrate with, extend, and build on, developers became an important user group in their own right. </span><strong><span>Developer experience (DX)</span></strong><span> grew out of the need to design for them specifically, with APIs, documentation, SDKs, CLIs, and workflows built around how developers work.</span></p><p><span>Now we&#8217;re building for a different kind of user: AI agents. Increasingly, agents are taking action on behalf of users and developers. They&#8217;re reading documentation, inspecting code, calling </span><a href="https://circleci.com/blog/what-is-api/"><span>APIs</span></a><span> and </span><a href="https://circleci.com/blog/mcp-vs-cli/"><span>CLIs</span></a><span>, and using </span><a href="https://circleci.com/product/mcp/"><span>MCP tools</span></a><span> to interact directly with products. And just as good UX and DX require designing around the needs of their users, products need to account for how agents find information, choose tools, make decisions, and take action.</span></p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!qTb-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a5aa789-e633-4b17-a80e-7fc308ef26fd_1840x940.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!qTb-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a5aa789-e633-4b17-a80e-7fc308ef26fd_1840x940.png 424w, https://substackcdn.com/image/fetch/$s_!qTb-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a5aa789-e633-4b17-a80e-7fc308ef26fd_1840x940.png 848w, https://substackcdn.com/image/fetch/$s_!qTb-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a5aa789-e633-4b17-a80e-7fc308ef26fd_1840x940.png 1272w, https://substackcdn.com/image/fetch/$s_!qTb-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a5aa789-e633-4b17-a80e-7fc308ef26fd_1840x940.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!qTb-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a5aa789-e633-4b17-a80e-7fc308ef26fd_1840x940.png" width="1456" height="744" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2a5aa789-e633-4b17-a80e-7fc308ef26fd_1840x940.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:744,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!qTb-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a5aa789-e633-4b17-a80e-7fc308ef26fd_1840x940.png 424w, https://substackcdn.com/image/fetch/$s_!qTb-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a5aa789-e633-4b17-a80e-7fc308ef26fd_1840x940.png 848w, https://substackcdn.com/image/fetch/$s_!qTb-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a5aa789-e633-4b17-a80e-7fc308ef26fd_1840x940.png 1272w, https://substackcdn.com/image/fetch/$s_!qTb-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2a5aa789-e633-4b17-a80e-7fc308ef26fd_1840x940.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The industry has started calling this </span><strong><a href="https://circleci.com/topics/agent-experience/"><span>agent experience (AX)</span></a></strong><span>. As more work moves through agents, their ability to understand and operate a product becomes part of the product experience itself. A platform can have excellent UX and DX and still be difficult for an agent to use reliably. </span><strong><span>At CircleCI, we&#8217;ve made solving that problem a core product goal: making CircleCI work just as well for agents as it does for the developers they support.</span></strong></p><p><span>In this issue we&#8217;re sharing what we&#8217;ve learned redesigning CircleCI for agents. We&#8217;ll start with the product problem that pushed us into AX, then walk through five findings from our testing that changed how we think about context, tools, evaluation, and the way agents interact with a product.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.confidentcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.confidentcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2><span>We started with broken builds</span></h2><p><span>The need to design for agents is already showing up in how customers use CircleCI. In our</span><a href="https://circleci.com/blog/five-takeaways-2026-software-delivery-report/"><span> 2026 software delivery data</span></a><span>, average daily workflow volume increased </span><strong><span>59% year over year</span></strong><span>, the largest increase we&#8217;ve measured. More agent-generated code means more builds, tests, and opportunities for failures to interrupt the development loop.</span></p><p><strong><span>Broken builds were the highest-value place to start.</span></strong><span> If agents can produce code quickly, they also need an efficient way to recover when validation fails. Our first approach followed what Netlify CEO Mathias Biilmann</span><a href="https://biilmann.blog/articles/introducing-ax/"><span> describes</span></a><span> as the </span><strong><span>closed</span></strong><span> model of agent experience. The original version of</span><a href="https://circleci.com/blog/introducing-chunk/"><span> </span></a><strong><a href="https://circleci.com/blog/introducing-chunk/"><span>Chunk</span></a></strong><span> was a hosted agent that could diagnose a failing build, propose a fix, and open a pull request.</span></p><p><span>In production, Chunk turned CI green on around half of its attempts, but developers chose to merge less than 25% of its fixes. The takeaway was clear: engineers wanted more control over the fixing process. So we moved toward the </span><strong><span>open</span></strong><span> model, giving their existing agents direct access to CircleCI data and tools, so the developer could stay informed and in control of the recovery process.</span></p><p><span>Early benchmark testing showed how sensitive agent performance was to the environment around the model. In a 100-sample benchmark of CI repair tasks, better log preprocessing increased the share of proposed fixes that closely matched the known solution from about </span><strong><span>25% to 53%</span></strong><span>. The size of the improvement made one thing clear: context, tools, evaluation, and workflow design could materially change how well the agent performed. We started testing those variables more deliberately, and five findings stood out.</span></p><h2><span>1. Agents need finely tuned context</span></h2><p><span>Our first instinct was to give the agent as much information as possible. The results moved in the opposite direction.</span></p><p><span>A raw build log can contain thousands of lines covering environment setup, dependency installation, test output, retries, warnings, and other execution details. Most of that information has little bearing on the failure the agent is trying to diagnose.</span></p><p><span>When we removed about 99% of the log content, the rate of fixes that took the wrong approach fell from 48% to 18%.</span></p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fzIg!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F938a1192-fc89-4375-9670-ad5067861aac_1840x840.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fzIg!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F938a1192-fc89-4375-9670-ad5067861aac_1840x840.png 424w, https://substackcdn.com/image/fetch/$s_!fzIg!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F938a1192-fc89-4375-9670-ad5067861aac_1840x840.png 848w, https://substackcdn.com/image/fetch/$s_!fzIg!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F938a1192-fc89-4375-9670-ad5067861aac_1840x840.png 1272w, https://substackcdn.com/image/fetch/$s_!fzIg!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F938a1192-fc89-4375-9670-ad5067861aac_1840x840.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fzIg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F938a1192-fc89-4375-9670-ad5067861aac_1840x840.png" width="1456" height="665" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/938a1192-fc89-4375-9670-ad5067861aac_1840x840.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:665,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!fzIg!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F938a1192-fc89-4375-9670-ad5067861aac_1840x840.png 424w, https://substackcdn.com/image/fetch/$s_!fzIg!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F938a1192-fc89-4375-9670-ad5067861aac_1840x840.png 848w, https://substackcdn.com/image/fetch/$s_!fzIg!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F938a1192-fc89-4375-9670-ad5067861aac_1840x840.png 1272w, https://substackcdn.com/image/fetch/$s_!fzIg!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F938a1192-fc89-4375-9670-ad5067861aac_1840x840.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The experiment also showed that context placement mattered. Adding git context as a file did not improve performance, but putting the same information directly into the message produced our strongest run at that point. The share of fixes that followed project conventions increased from 59% to 90%.</span></p><p><span>For teams building agent workflows, context is worth treating as an input you design and evaluate. Test what information you provide, how much you provide, and where the agent encounters it.</span></p><h2><span>2. Keep the tool set focused</span></h2><p><span>Better context improved performance, so we tested whether giving the agent more capabilities would help further. We added skills, forced reasoning steps, and a much larger set of </span><strong><span>MCP tools</span></strong><span>.</span></p><p><span>More capability did not consistently produce better results. Added skills used about 24% more tokens without improving quality, and forced workflow steps reduced fix rates in our tests. In our largest MCP test, the server exposed 138 tools for the agent to choose from. The agent consistently relied on just 10 of them, less than 10% of the available tool set.</span></p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!UtgZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aefdfe-bf60-4366-87e1-c0f8e80b78c7_1840x720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!UtgZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aefdfe-bf60-4366-87e1-c0f8e80b78c7_1840x720.png 424w, https://substackcdn.com/image/fetch/$s_!UtgZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aefdfe-bf60-4366-87e1-c0f8e80b78c7_1840x720.png 848w, https://substackcdn.com/image/fetch/$s_!UtgZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aefdfe-bf60-4366-87e1-c0f8e80b78c7_1840x720.png 1272w, https://substackcdn.com/image/fetch/$s_!UtgZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aefdfe-bf60-4366-87e1-c0f8e80b78c7_1840x720.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!UtgZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aefdfe-bf60-4366-87e1-c0f8e80b78c7_1840x720.png" width="1456" height="570" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a2aefdfe-bf60-4366-87e1-c0f8e80b78c7_1840x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:570,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!UtgZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aefdfe-bf60-4366-87e1-c0f8e80b78c7_1840x720.png 424w, https://substackcdn.com/image/fetch/$s_!UtgZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aefdfe-bf60-4366-87e1-c0f8e80b78c7_1840x720.png 848w, https://substackcdn.com/image/fetch/$s_!UtgZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aefdfe-bf60-4366-87e1-c0f8e80b78c7_1840x720.png 1272w, https://substackcdn.com/image/fetch/$s_!UtgZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa2aefdfe-bf60-4366-87e1-c0f8e80b78c7_1840x720.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>A developer who knows a product can navigate a broad API surface and choose the right endpoint. An agent has to select among MCP tools based on their names, descriptions, and the context available to it. Exposing every underlying API capability as a separate tool can make that choice harder without helping the agent complete the task.</span></p><p><span>For teams designing MCP servers and other agent interfaces, start with the actions required to complete the job. Add tools when testing shows that the additional capability improves performance.</span></p><h2><span>3. Evaluate the result you want to ship</span></h2><p><span>One of our early evaluation criteria was simple: did the agent make the build pass?</span></p><p><span>A passing build turned out to be an incomplete measure. An agent can turn CI green by increasing a timeout, loosening an assertion, excluding a failing package, or changing the failing test instead of fixing the code that caused the failure.</span></p><p><span>We started measuring </span><strong><span>merge quality</span></strong><span> separately from </span><strong><span>fix rate</span></strong><span>: not only whether CI passed, but whether the proposed fix was good enough to merge.</span></p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lKzD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624f46e9-5c7d-43c6-86ca-d76e423bb8d0_1840x1120.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lKzD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624f46e9-5c7d-43c6-86ca-d76e423bb8d0_1840x1120.png 424w, https://substackcdn.com/image/fetch/$s_!lKzD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624f46e9-5c7d-43c6-86ca-d76e423bb8d0_1840x1120.png 848w, https://substackcdn.com/image/fetch/$s_!lKzD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624f46e9-5c7d-43c6-86ca-d76e423bb8d0_1840x1120.png 1272w, https://substackcdn.com/image/fetch/$s_!lKzD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624f46e9-5c7d-43c6-86ca-d76e423bb8d0_1840x1120.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lKzD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624f46e9-5c7d-43c6-86ca-d76e423bb8d0_1840x1120.png" width="1456" height="886" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/624f46e9-5c7d-43c6-86ca-d76e423bb8d0_1840x1120.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:886,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!lKzD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624f46e9-5c7d-43c6-86ca-d76e423bb8d0_1840x1120.png 424w, https://substackcdn.com/image/fetch/$s_!lKzD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624f46e9-5c7d-43c6-86ca-d76e423bb8d0_1840x1120.png 848w, https://substackcdn.com/image/fetch/$s_!lKzD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624f46e9-5c7d-43c6-86ca-d76e423bb8d0_1840x1120.png 1272w, https://substackcdn.com/image/fetch/$s_!lKzD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F624f46e9-5c7d-43c6-86ca-d76e423bb8d0_1840x1120.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The benchmark exposed a meaningful gap: about 90% of fixes made CI pass, but only about 77% met our merge-quality bar. A green build alone was not always a sufficient measure of success.</span></p><p><span>The same evaluation problem applies to other agent systems. Intermediate metrics are useful, but the final evaluation needs to represent the outcome the user would accept.</span></p><h2><span>4. Test the same task more than once</span></h2><p><span>Changing the evaluation criteria also exposed variation between repeated runs.</span></p><p><span>The agent usually diagnosed the problem correctly, but the same failure did not always produce the same quality of fix. Incomplete changes and fixes that expanded beyond the required scope accounted for much of the variation.</span></p><p><span>Across repeated benchmark runs, </span><strong><span>27% of cases flipped between pass and fail</span></strong><span>.</span></p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!5wHM!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ff4bfb1-cf02-489d-96c1-c9f6c500452a_1840x840.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!5wHM!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ff4bfb1-cf02-489d-96c1-c9f6c500452a_1840x840.png 424w, https://substackcdn.com/image/fetch/$s_!5wHM!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ff4bfb1-cf02-489d-96c1-c9f6c500452a_1840x840.png 848w, https://substackcdn.com/image/fetch/$s_!5wHM!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ff4bfb1-cf02-489d-96c1-c9f6c500452a_1840x840.png 1272w, https://substackcdn.com/image/fetch/$s_!5wHM!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ff4bfb1-cf02-489d-96c1-c9f6c500452a_1840x840.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!5wHM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ff4bfb1-cf02-489d-96c1-c9f6c500452a_1840x840.png" width="1456" height="665" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2ff4bfb1-cf02-489d-96c1-c9f6c500452a_1840x840.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:665,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!5wHM!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ff4bfb1-cf02-489d-96c1-c9f6c500452a_1840x840.png 424w, https://substackcdn.com/image/fetch/$s_!5wHM!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ff4bfb1-cf02-489d-96c1-c9f6c500452a_1840x840.png 848w, https://substackcdn.com/image/fetch/$s_!5wHM!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ff4bfb1-cf02-489d-96c1-c9f6c500452a_1840x840.png 1272w, https://substackcdn.com/image/fetch/$s_!5wHM!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ff4bfb1-cf02-489d-96c1-c9f6c500452a_1840x840.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Clearer task boundaries helped reduce some of the variation. When the requested scope was narrow and explicit, the agent had less room to make unrelated changes or expand the solution beyond the failure.</span></p><p><span>For teams evaluating agent workflows, repeat the same tasks and look at the distribution of outcomes. A strong average can hide variation that matters when the system is used repeatedly.</span></p><h2><span>5. Give agents direct access to product data</span></h2><p><span>Cleaner inputs helped once the agent had the information it needed. We also found that agents spent a significant part of each run obtaining and moving information before they could work on the fix.</span></p><p><span>CircleCI already has much of what an agent needs to begin: which workflow failed, which job produced the error, which commit was running, what tests failed, and what happened during execution.</span></p><p><span>In our testing, only </span><strong><span>20&#8211;30% of a run</span></strong><span> was spent reading the error and editing code. The rest included surrounding work such as exploring the repository, working with git, and moving information between systems.</span></p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!S7RP!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf18d3df-b021-45fb-8ce3-730a10cc8015_1840x1580.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!S7RP!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf18d3df-b021-45fb-8ce3-730a10cc8015_1840x1580.png 424w, https://substackcdn.com/image/fetch/$s_!S7RP!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf18d3df-b021-45fb-8ce3-730a10cc8015_1840x1580.png 848w, https://substackcdn.com/image/fetch/$s_!S7RP!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf18d3df-b021-45fb-8ce3-730a10cc8015_1840x1580.png 1272w, https://substackcdn.com/image/fetch/$s_!S7RP!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf18d3df-b021-45fb-8ce3-730a10cc8015_1840x1580.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!S7RP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf18d3df-b021-45fb-8ce3-730a10cc8015_1840x1580.png" width="1456" height="1250" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/cf18d3df-b021-45fb-8ce3-730a10cc8015_1840x1580.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1250,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!S7RP!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf18d3df-b021-45fb-8ce3-730a10cc8015_1840x1580.png 424w, https://substackcdn.com/image/fetch/$s_!S7RP!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf18d3df-b021-45fb-8ce3-730a10cc8015_1840x1580.png 848w, https://substackcdn.com/image/fetch/$s_!S7RP!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf18d3df-b021-45fb-8ce3-730a10cc8015_1840x1580.png 1272w, https://substackcdn.com/image/fetch/$s_!S7RP!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fcf18d3df-b021-45fb-8ce3-730a10cc8015_1840x1580.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Longer feedback loops also added measurable cost. In one iterative experiment, runtime increased from about </span><strong><span>10 to 20 minutes</span></strong><span> and token usage from </span><strong><span>2.4M to 3.6M</span></strong><span> as the agent kept working toward a passing result.</span></p><p><span>For teams adding agent access to an existing product, </span><strong><span>look closely at the information your product already has when an agent begins a task</span></strong><span>. Direct access can reduce the searching and reconstruction the agent would otherwise need to perform itself.</span></p><h2><span>How we&#8217;re improving agent experience at CircleCI</span></h2><p><span>Across the experiments, we eventually reached a point where a capable model could turn CI green on roughly </span><strong><span>nine in ten benchmark cases</span></strong><span>: </span><strong><span>90% in one run and 84% in the repeat</span></strong><span>.</span></p><p><span>Reaching that level required substantial experimentation with context, tools, evaluation, and feedback loops. About </span><strong><span>77% of benchmark fixes met our merge-quality bar</span></strong><span>, but production usage showed that developers wanted to stay closer to the fixing process, and the share of fixes they actually merged remained much lower than we wanted.</span></p><p><span>Those findings shaped the AX work we&#8217;re shipping now. We&#8217;re taking what made the fixer effective and building it into interfaces that give developers&#8217; chosen agents first-class access to CircleCI data and actions.</span></p><p><span>A lot of the infrastructure behind the new experience is new or substantially rewritten, and we&#8217;re rolling out improvements on a daily cadence.</span></p><p><strong><span>Recent changes include:</span></strong></p><ul><li><p><span>A </span><a href="https://circleci.com/blog/rebuilding-the-circleci-cli-from-scratch/"><span>ground-up rewrite of the </span></a><strong><a href="https://circleci.com/blog/rebuilding-the-circleci-cli-from-scratch/"><span>CircleCI CLI</span></a></strong><span> for developers and agents, with predictable JSON, stable exit behavior, agent-readable errors, and a built-in MCP server.</span></p></li><li><p><span>circleci run get --failure-report, which gives agents condensed failure diagnostics instead of requiring them to work through complete build logs.</span></p></li><li><p><strong><span>Copy Fix Prompt</span></strong><span>, which packages the relevant CircleCI failure context into a prompt developers can paste into the coding agent they already use.</span></p></li><li><p><span>Expanded </span><strong><span>MCP tools</span></strong><span> for inspecting runs, jobs, build output, and test results and taking CI actions from an agent&#8217;s development environment.</span></p></li><li><p><strong><span>OAuth 2.0 with Dynamic Client Registration and PKCE</span></strong><span>, giving local tools and agent integrations a standard way to request CircleCI API access without a manually provisioned client secret.</span></p></li></ul><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!sn-P!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd98661bb-5b10-4f65-a973-bd4c7c642a60_2048x804.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!sn-P!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd98661bb-5b10-4f65-a973-bd4c7c642a60_2048x804.png 424w, https://substackcdn.com/image/fetch/$s_!sn-P!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd98661bb-5b10-4f65-a973-bd4c7c642a60_2048x804.png 848w, https://substackcdn.com/image/fetch/$s_!sn-P!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd98661bb-5b10-4f65-a973-bd4c7c642a60_2048x804.png 1272w, https://substackcdn.com/image/fetch/$s_!sn-P!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd98661bb-5b10-4f65-a973-bd4c7c642a60_2048x804.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!sn-P!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd98661bb-5b10-4f65-a973-bd4c7c642a60_2048x804.png" width="1456" height="572" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d98661bb-5b10-4f65-a973-bd4c7c642a60_2048x804.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:572,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!sn-P!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd98661bb-5b10-4f65-a973-bd4c7c642a60_2048x804.png 424w, https://substackcdn.com/image/fetch/$s_!sn-P!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd98661bb-5b10-4f65-a973-bd4c7c642a60_2048x804.png 848w, https://substackcdn.com/image/fetch/$s_!sn-P!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd98661bb-5b10-4f65-a973-bd4c7c642a60_2048x804.png 1272w, https://substackcdn.com/image/fetch/$s_!sn-P!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd98661bb-5b10-4f65-a973-bd4c7c642a60_2048x804.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><em><span>The new Copy Fix Prompt option is one of the ways we&#8217;re making CircleCI more friendly to agentic workflows.</span></em></p><p><span>The details are still moving as we put these interfaces in front of more real agent workflows. Our goal is to improve both sides of the experience: help agents diagnose and fix CI problems well, while giving developers the visibility and control they need to work with those fixes.</span></p><p><span>You can try the current fixing flow from any failing CircleCI workflow:</span></p><ul><li><p><span>Install the new CircleCI CLI: </span><code>brew install circleci</code></p></li><li><p><span>Run </span><code>circleci run get --failure-report</code><span> on a failing run to get condensed diagnostics for your local agent</span></p></li><li><p><span>Or use </span><strong><span>Copy Fix Prompt</span></strong><span> from a failing workflow for a manual handoff</span></p></li><li><p><strong><span>Need to sign up?</span></strong><span> Run </span><code>circleci onboard</code><span> from the CLI to get started, or</span><a href="https://app.circleci.com/signup"><span> create a free CircleCI account from our signup page</span></a><span>.</span></p></li></ul><p><span>The biggest lesson from our testing is that agent experience is shaped by far more than the model. Context, tool design, evaluation, repeatability, and direct access to product data all changed how reliably agents could work with CircleCI. For teams building for agents, AX means designing those surrounding systems as deliberately as you design the product experience for people.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.confidentcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Bake low Merge Efficiency Ratio (MER) into the agent loop]]></title><description><![CDATA[How engineering your agent loop for low MER (with inner-loop checks, a CI sneak peek, and a single push) cuts outer-loop CI credits by 5&#8211;6x without adding time or cost.]]></description><link>https://www.confidentcommit.com/p/bake-low-merge-efficiency-ratio-mer</link><guid isPermaLink="false">https://www.confidentcommit.com/p/bake-low-merge-efficiency-ratio-mer</guid><dc:creator><![CDATA[Ryan E. Hamilton]]></dc:creator><pubDate>Tue, 22 Sep 2026 15:01:16 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b63ee9b1-852e-4892-a1da-d9a560c28bdf_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>CircleCI&#8217;s </span><a href="https://circleci.com/blog/five-takeaways-2026-q2-pulse/"><span>2026 State of Software Delivery Q2 Pulse Report</span></a><span> named the key performance metric we&#8217;re talking about here: </span><strong><span>Merge Efficiency Ratio (MER)</span></strong><span>. How many feature-branch validation cycles does it take to get a change onto main?</span></p><p><span>Median teams sit around </span><strong><span>3.9</span></strong><span>. The top 5% run about </span><strong><span>2.6</span></strong><span>. An elite cohort of twenty orgs is already near </span><strong><span>1.3</span></strong><span>.</span></p><p><span>Most of the industry is still grinding through several rounds of feature-branch rework before earning the right to merge to main. The leaders are closer to one.</span></p><p><span>When coding agents are the ones waiting on those rework cycles, MER is not a vanity metric. It is the clock and the bill. Extra loops mean more tokens burned, more CI credits spent, and more agent context rotting on a red light from CI.</span></p><p><span>The Q2 Pulse Report is blunt: </span><strong><span>lower MER is a DX win, an AX win, and a cost win.</span></strong></p><p><span>I ran a small AFK lab experiment that rhymes with that report. Not an org-level MER study. Same shape of problem, inside a single agent setup.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.confidentcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.confidentcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2><span>Hypothesis</span></h2><p><span>If a green PR is already likely, pushing only once at the end should burn far fewer outer-loop CI credits than pushing after every task. And if deterministic checks live on the inner loop </span><strong><span>once</span></strong><span>, you stop paying time and money three times for the same linter.</span></p><p><span>Pushing once at the end is only safe once the agent can reliably go green.</span></p><h2><span>Setup</span></h2><p><span>In our core experiment to determine </span><a href="https://www.confidentcommit.com/p/cost-of-a-green-pr"><span>the cost of a green PR</span></a><span>, we found that if you give the agent a sneak peek at the CI setup, let it take notes before coding begins, and let it do a practice run on inner-loop checks before pushing to outer-loop CI, you can achieve 100% green PRs with zero rework.</span></p><p><span>The sneak peek is actually quite simple. Before task one, before a line of game code gets written, the agent receives a short inventory of the CI setup: which checks fire on the inner-loop practice run, which jobs fire on the outer-loop CI pipeline, and pointers to the scripts and configs that define them. This information is descriptive, but not an answer key. The agent writes its own notes from that inventory </span><code>preflight.md</code><span>) and can </span><code>@</code><span>-read those configs whenever it wants more. In our codebase, the inventory is a &#8220;CI validation manifest,&#8221; which tells the agent what will be graded and where the grading logic lives, before it starts guessing from vibes.</span></p><p><strong><span>Fix loops drop to zero. Green PRs go to 100%.</span></strong></p><p><span>This piece starts there. Same </span><a href="https://www.confidentcommit.com/p/what-snake-games-have-taught-us-about"><span>AFK Snake game build</span></a><span>. Same seven tasks. Same two cadences, now with the sneak peek on.</span></p><ul><li><p><strong><span>Per-task-push:</span></strong><span> a pipeline after every task. Seven tasks, seven trips.</span></p></li><li><p><strong><span>Single-push:</span></strong><span> finish the stack, clear practice run, push once at the end.</span></p></li></ul><p><span>How often you push is the lever under test. First you have to be equipped to push green.</span></p><p><span>A linter is a linter. Compute is compute, whether it runs on a laptop, in a </span><a href="https://circleci.com/blog/chunk-sidecars/"><span>Chunk sidecar</span></a><span> microbuild, or in a cloud job financed with CI credits.</span></p><p><span>There is no real reason to run the same deterministic check on localhost, again on the sidecar, and again on the outer-loop CI pipeline, as if three identical stamps make the code more correct.</span></p><p><span>That stack is a relic of human by-hand engineering. A person in flow forgets to run the test suite, pushes red, and we paper over forgetfulness by re-running the universe on every layer. Coding agents do not forget the same way. You can (and should) instruct them to clear inner-loop checks </span><strong><span>before</span></strong><span> they spend a single outer-loop CI credit.</span></p><p><span>In this calibration we moved inner-loop checks fully onto the sidecar and left thick outer-loop CI as the honest final test. Slimming true duplicates on the outer-loop surface is later work. The principle does not wait though: </span><strong><span>do not triple-pay for the same answer.</span></strong></p><h2><span>Results</span></h2><p><span>Headline: once sneak peek made 100% green likely, a single end-of-run push cut estimated CI credits from </span><strong><span>~81 to ~14</span></strong><span>.</span></p><ul><li><p><strong><span>Per-task-push:</span></strong><span> </span><strong><span>7</span></strong><span> pipelines. </span><strong><span>~81.4</span></strong><span> estimated CI credits.</span></p></li><li><p><strong><span>Single-push:</span></strong><span> </span><strong><span>1</span></strong><span> pipeline. </span><strong><span>~14.0</span></strong><span> estimated credits.</span></p></li></ul><p><strong><span>About 5 to 6x fewer CI credits</span></strong><span> on the single-push run.</span></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HKy-!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed489ada-703a-40d0-b076-9b12304c3384_1746x418.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HKy-!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed489ada-703a-40d0-b076-9b12304c3384_1746x418.png 424w, https://substackcdn.com/image/fetch/$s_!HKy-!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed489ada-703a-40d0-b076-9b12304c3384_1746x418.png 848w, https://substackcdn.com/image/fetch/$s_!HKy-!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed489ada-703a-40d0-b076-9b12304c3384_1746x418.png 1272w, https://substackcdn.com/image/fetch/$s_!HKy-!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed489ada-703a-40d0-b076-9b12304c3384_1746x418.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HKy-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed489ada-703a-40d0-b076-9b12304c3384_1746x418.png" width="1456" height="349" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ed489ada-703a-40d0-b076-9b12304c3384_1746x418.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:349,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:94613,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/216368190?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed489ada-703a-40d0-b076-9b12304c3384_1746x418.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HKy-!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed489ada-703a-40d0-b076-9b12304c3384_1746x418.png 424w, https://substackcdn.com/image/fetch/$s_!HKy-!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed489ada-703a-40d0-b076-9b12304c3384_1746x418.png 848w, https://substackcdn.com/image/fetch/$s_!HKy-!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed489ada-703a-40d0-b076-9b12304c3384_1746x418.png 1272w, https://substackcdn.com/image/fetch/$s_!HKy-!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fed489ada-703a-40d0-b076-9b12304c3384_1746x418.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><span>Wall-clock stayed in the </span><strong><span>50 to 60 minute</span></strong><span> band either way (55.5 vs 56.7). LLM dollars stayed in the </span><strong><span>$15 to $17</span></strong><span> band (17.15 vs 15.73). Tokens still track agent work. The CI credit collapse comes from the push cadence, not a cheaper model.</span></p><p><span>That bargain only holds when a sneak peek plus inner-loop checks make a 100% green PR likely. If you pile up unproven work and trigger the outer-loop pipeline once cold, you have not saved anything. You have postponed finding out you are red. When that one pipeline fails, you still pay the fix loop, and a retry can pick up new last-mile surprises.</span></p><p><strong><span>Earn one take. Then trigger outer-loop CI once.</span></strong></p><h2><span>TL;DR</span></h2><p><strong><span>Elite engineering teams do not win by loving rework. They win by needing less of it.</span></strong></p><p><span>Think of agent-era validation as a continuum. Inner loop to outer loop, hybrid on purpose. Each check earns its place.</span></p><ul><li><p><strong><span>Deterministic checks live inward:</span></strong><span> cheap, early, once.</span></p></li><li><p><strong><span>World-shaped risk lives outward:</span></strong><span> policy, last-mile, flaky integrations, the CVE that did not exist at 9am. You name it.</span></p></li></ul><p><span>Place each check where its information is unique. Everything else is just another duplicate cost for the same answer.</span></p><p><span>Outer-loop CI should still exist. It is where the world gets a vote. It should not be a museum of jobs you already passed in identical form on the previous two layers of validation.</span></p><p><span>At AFK scale, the inheritance path looks like this:</span></p><ol><li><p><strong><span>Make preventable reds disappear</span></strong><span> (</span><a href="https://www.confidentcommit.com/p/cost-of-a-green-pr"><span>push a /green PR in one take</span></a><span>).</span></p></li><li><p><strong><span>Keep remaining retries cheap</span></strong><span> (use the </span><code>--failure-report</code><span> flag).</span></p></li><li><p><strong><span>Put deterministic checks on the inner loop once.</span></strong><span> Push to outer-loop CI as often as your green rate deserves.</span></p></li></ol><p><span>Once 100% green is something you can count on, pushing only once at the end of the run is an obvious CI credit win.</span></p><p><span>Bake that into your agent setup. Low MER should not be a hero dashboard you inspect after the damage. It should be the default the agent loop was engineered to produce.</span></p><p><span>Rework should be reserved for unknown failures from the real world. Preventable failures should be totally eliminated.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.confidentcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div>]]></content:encoded></item><item><title><![CDATA[The only context LLMs can’t fake]]></title><description><![CDATA[Rob Zuber sits down with Dennis Pilarinos to talk about what agents don&#8217;t know and why it&#8217;s slowing your team down.]]></description><link>https://www.confidentcommit.com/p/your-agent-is-a-day-one-engineer</link><guid isPermaLink="false">https://www.confidentcommit.com/p/your-agent-is-a-day-one-engineer</guid><dc:creator><![CDATA[Confident Commit]]></dc:creator><pubDate>Tue, 15 Sep 2026 20:23:26 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/215874289/492246d07a5134d38c0cbc082c991e21.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><span>There&#8217;s a gap that most engineering teams are starting to feel but haven&#8217;t quite named yet. You spin up an agent, give it a task, and it goes... somewhere. It reads files, checks documentation, wanders down paths that don&#8217;t pan out. It&#8217;s not incompetent, it just doesn&#8217;t know how things work </span><em><span>here</span></em><span>.</span></p><p><span>Dennis Pilarinos has been thinking about this problem for years. He&#8217;s the founder and CEO of Unblocked, which he describes as the context layer for modern engineering teams: a system that surfaces the tribal knowledge inside an organization and makes it available to both the people doing the work and the agents working alongside them. Before Unblocked, Dennis spent years at Microsoft on the early Azure team, then moved through mobile CI and developer tooling at a time when, as he puts it, investors were telling founders &#8220;there&#8217;s no money in DevTools.&#8221;</span></p><p><span>There&#8217;s money in DevTools now. But the more interesting question is whether teams are building the right things into them. Rob and Dennis went deep on that question: why agents fail in ways humans don&#8217;t, what happens to the pull request when agents are writing the code, and whether we&#8217;re finally close to being able to draw a line from &#8220;we ran this agent&#8221; to &#8220;we moved this metric.&#8221;</span></p><h2><span>1. Every agent you spin up starts on day one. Every time.</span></h2><p><span>The framing Dennis uses here is hard to shake: imagine joining a company and being asked to fix a bug. You&#8217;re reasonably competent. But you&#8217;re up against someone who has three to seven years of context on that codebase. Who gets it done better? Almost certainly not you, not yet.</span></p><p><span>That&#8217;s the situation for every agent, every session. It doesn&#8217;t accumulate experience across runs. It doesn&#8217;t remember the incident channel discussion from three months ago, or why that particular API pattern was chosen over the obvious alternative. Each session starts cold.</span></p><blockquote><p><span>&#8220;When you spin up an agent, it is basically like me joining CircleCI on day one. It has no meaningful historical context for which to draw upon. And so it does what I would do. I would probably dig through a whole bunch of documentation. I would go down paths that don&#8217;t actually lead to productive outcomes.&#8221;</span></p><p><strong><span>Dennis Pilarinos, Founder and CEO, Unblocked</span></strong></p></blockquote><p><span>The practical takeaway is this: if your agents are producing messy, roundabout solutions, the problem is likely the missing context. What the agent needs is often sitting in Slack, in a closed issue, in a Confluence doc nobody&#8217;s opened in 18 months.</span></p><h2><span>2. The code alone is not enough context.</span></h2><p><span>Dennis is direct about something most teams discover the hard way: indexing the codebase is a starting point, not a solution. The reasoning behind the code lives elsewhere.</span></p><blockquote><p><span>&#8220;I have never ever met a company that&#8217;s like, our docs are super well organized and completely up to date. That just doesn&#8217;t exist.&#8221;</span></p><p><strong><span>Dennis Pilarinos, Founder and CEO, Unblocked</span></strong></p></blockquote><p><span>When Unblocked pulls in Slack threads, bug trackers, incident channels, and documentation alongside source code, the quality of outcomes goes up significantly. This isn&#8217;t a surprise in retrospect. It&#8217;s exactly how a person learns a codebase. They read pull request comments. They get feedback on their first few contributions. They sit in on postmortems. The &#8220;scar tissue,&#8221; as Dennis calls it, is where the real knowledge lives.</span></p><p><span>Teams building context systems that rely only on source code are building the equivalent of onboarding a new engineer with read-only access to the repo and no other communication. It doesn&#8217;t work for people either.</span></p><h2><span>3. AI code review is only as good as the context behind it.</span></h2><p><span>Unblocked started using its own context layer to build an internal code review tool after surveying the existing market and finding it underwhelming. The tools they evaluated felt like early-stage engineers leaving noise in PRs, sometimes literally leaving haikus. Not useful.</span></p><p><span>When they built review on top of a context layer that knew the history of decisions, past incidents, and team-specific conventions, the results were different. The reviewer understood </span><em><span>why</span></em><span> certain patterns existed.</span></p><blockquote><p><span>&#8220;It&#8217;s not just limited to the source code. It knows conversations you&#8217;ve had about incidents, or some of the best practices that you might have written down somewhere else that other tools might not necessarily know.&#8221;</span></p><p><strong><span>Dennis Pilarinos, Founder and CEO, Unblocked</span></strong></p></blockquote><p><span>If you&#8217;re evaluating AI code review tools and finding them shallow, this is worth paying attention to. The limiting factor is usually not the model&#8217;s ability to read code. It&#8217;s whether the model understands the context in which that code was written. Most tools don&#8217;t have that. The ones that will matter will.</span></p><h2><span>4. The pull request is two tools in a trench coat.</span></h2><p><span>Rob and Dennis had one of the more useful framings of this conversation on the subject of pull requests. The PR solves two distinct problems: it&#8217;s an information-sharing mechanism (here&#8217;s what changed and why) and a change management mechanism (here&#8217;s how we control what goes to production). Those are different problems, and they&#8217;ve been bundled together for so long that most teams have stopped asking whether they need to be.</span></p><p><span>As agents generate more code, the PR review bottleneck gets real fast. Teams that Dennis talks to are saying their engineers are backed up on review, not on writing. The question becomes: can you decompose what the PR was doing and solve each part better?</span></p><blockquote><p><span>&#8220;Go back to first principles. Pull that apart and say, what would be the way we would solve for this piece of it, for this specific problem in this new context? It doesn&#8217;t have to be one solution. It just happens to be now, because it was there.&#8221; </span><strong><span>Rob Zuber, CTO, CircleCI</span></strong></p></blockquote><p><span>The implication for engineering leaders: look at every gate in your delivery pipeline and ask what problem it was originally solving. The answer might still be worth solving. The mechanism might not be the right one anymore.</span></p><h2><span>5. AI adoption inside one company spans an enormous range.</span></h2><p><span>This was one of the more grounding moments in the conversation. Dennis described organizations where some engineers are just discovering tab-completion and others are running thousands of autonomous agents handling multi-hour implementation tasks. Same company. Same week.</span></p><p><span>He puts roughly three to five percent of engineering teams in the &#8220;AI pilled&#8221; category: people who have experienced enough of the technology to feel acutely where the gaps are. They&#8217;re the ones who find the lack of organizational context genuinely painful, because they&#8217;ve hit the ceiling. The rest of the org might not have reached that ceiling yet.</span></p><p><span>This matters for how you roll out tooling and expectations. Someone whose primary interaction with AI is autocomplete doesn&#8217;t have the same mental model of what&#8217;s possible, or of what&#8217;s broken. The journey isn&#8217;t linear, and it&#8217;s not uniform.</span></p><h2><span>6. The ROI problem is closer to being solvable than it&#8217;s ever been.</span></h2><p><span>This was the most forward-looking thread in the conversation, and both Rob and Dennis were careful about timelines. But the observation is worth sitting with: agents are instrumented in ways that people never were. You can actually see what it cost to get to a solution. The question is whether you can connect that to business outcomes.</span></p><blockquote><p><span>&#8220;We can say: what did it take in order for you to come to this conclusion or to this solution? And can you do it in a token-efficient way? But what will happen next is: how do we reconcile that execution relative to business objectives? Not just how much did it cost, but how did it actually move the needle for the business?&#8221;</span></p><p><strong><span>Dennis Pilarinos, Founder and CEO, Unblocked</span></strong></p></blockquote><p><span>Finance teams are noticing the spike in costs. The engineering side hasn&#8217;t yet built the case for what that cost produced. Closing that gap, drawing the through-line from token spend to business outcome, is the work in front of engineering leaders right now. Dennis thinks we&#8217;re closer than we&#8217;ve ever been. Rob&#8217;s take: we&#8217;ve taken one step on a very long journey, but it&#8217;s a real step.</span></p><p><em><span>Confident Commit is CircleCI&#8217;s publication for engineering leaders and practitioners. Subscribe to get episodes and companion posts delivered directly.</span></em></p>]]></content:encoded></item><item><title><![CDATA[We built a new API for AI agents. None of them used it (directly). ]]></title><description><![CDATA[Agents make 97% of their calls against our new agent-friendly API. Not one agent found the API by reading the catalog, the llms.txt or the OpenAPI spec.]]></description><link>https://www.confidentcommit.com/p/we-built-a-new-api-for-ai-agents</link><guid isPermaLink="false">https://www.confidentcommit.com/p/we-built-a-new-api-for-ai-agents</guid><dc:creator><![CDATA[Dan Mullineux]]></dc:creator><pubDate>Tue, 15 Sep 2026 16:11:41 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/bec5e621-0c78-48a5-8808-0aa400b2843f_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Every company with an API has been busy making it AI friendly, with the goal of improving how agents interact with their products. That is part of the advice that big model providers </span><a href="https://www.anthropic.com/engineering/writing-tools-for-agents"><span>suggest</span></a><span>.</span></p><p><span>Our current CircleCI APIs use version numbers in the URLs. V1, V1.1 and V2. They had accumulated four incompatible URL patterns, three ID formats and a different response shape per endpoint, effectively ten years of drift. The pitch for a new API V3 was explicitly about consistency for agents: one URL grammar, UUIDs everywhere, one response envelope, one error shape. An LLM should be able to predict </span><code>/api/v3/{entity}/:id</code><span> from the entity name alone instead of carrying the whole spec in its context.</span></p><p><span>We followed all the current advice on how to make the API discoverable by agents. We published the OpenAPI spec at a stable URL. We added </span><code>/.well-known/api-catalog per RFC 9727</code><span>, so a machine could discover the machine-readable descriptions. We generated an </span><code>llms.txt</code><span> and an </span><code>llms-full.txt</code><span>. We rendered the whole reference as markdown as well as HTML. We put the conventions - error shapes, idempotency, which operations are irreversible - into the spec description so they travel with every entity document, not just the index.</span></p><p><span>This post answers one question: did any of that work?</span></p><p></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!c3Z5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1376e541-4709-4d04-9a8d-eb61f3b148d7_1956x420.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!c3Z5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1376e541-4709-4d04-9a8d-eb61f3b148d7_1956x420.png 424w, https://substackcdn.com/image/fetch/$s_!c3Z5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1376e541-4709-4d04-9a8d-eb61f3b148d7_1956x420.png 848w, https://substackcdn.com/image/fetch/$s_!c3Z5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1376e541-4709-4d04-9a8d-eb61f3b148d7_1956x420.png 1272w, https://substackcdn.com/image/fetch/$s_!c3Z5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1376e541-4709-4d04-9a8d-eb61f3b148d7_1956x420.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!c3Z5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1376e541-4709-4d04-9a8d-eb61f3b148d7_1956x420.png" width="1456" height="313" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1376e541-4709-4d04-9a8d-eb61f3b148d7_1956x420.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:313,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!c3Z5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1376e541-4709-4d04-9a8d-eb61f3b148d7_1956x420.png 424w, https://substackcdn.com/image/fetch/$s_!c3Z5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1376e541-4709-4d04-9a8d-eb61f3b148d7_1956x420.png 848w, https://substackcdn.com/image/fetch/$s_!c3Z5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1376e541-4709-4d04-9a8d-eb61f3b148d7_1956x420.png 1272w, https://substackcdn.com/image/fetch/$s_!c3Z5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1376e541-4709-4d04-9a8d-eb61f3b148d7_1956x420.png 1456w" sizes="100vw" fetchpriority="high"></picture><div></div></div></a></figure></div><h2><span>Hypothesis</span></h2><p><span>My assumption at the outset was that agents would </span><em><span>want to work with</span></em><span> a consistent API, and that they would be able to easily find it.</span></p><p><span>I expected consistency to be the hard part and discovery to be the easy win.</span></p><p><span>Specifically:</span></p><ol><li><p><span>Agents would find V3 on their own through the discovery surface - the catalog, the OpenAPI spec, </span><code>llms.txt</code><span>.</span></p></li><li><p><span>Once they found it, V3&#8217;s consistency would keep them there, because guessing a V3 URL is easier than looking one up in V2.</span></p></li><li><p><span>V2 usage by agents would decline as agents discovered the better option.</span></p></li></ol><p><span>The first and third were wrong in ways more interesting than being right.</span></p><h2><span>Setup</span></h2><p><span>I first proposed building a new consistent API 3 years ago. At the time, it would have been a massive time expenditure likely not worth the investment. But by the time I pitched it again earlier this year, the execution had changed; we could use a swarm of agents with good skills. Agentic workflow enabled us to decide that it was a low enough cost to take on an experiment previously considered to be very high cost, and impacting every team on the org. I got the green light.</span></p><p><span>I tend to be more pragmatic than scientific. I understand the scientific method, but can also plan ahead to mitigate the risks of just diving in and trying things. In the case of the API, although it&#8217;s fairly high risk in some ways, I knew that if we didn&#8217;t publicize it fully, we could always stop and roll back.</span></p><p><span>Every V3 route ingresses through one gateway service. That was a deliberate call made for enforcement reasons, and it is also what makes this measurable: one dataset holds every public API request with its route, method and user agent.</span></p><p><span>After agreeing on the consistent flexible API shape, we built skills internally for teams to use to guide their agents, and within a few months we had V3 APIs in place that could replace 90% of existing API traffic.</span></p><p><span>Measuring agent traffic hitting our API is an inexact science, relying on parsing the User-Agent in the inbound requests.</span></p><p><span>Real strings from production for requests coming in via our CLI:</span></p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;dcb7ceee-ffba-41da-ad37-761c3ac94319&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">circleci-cli (darwin/arm64; 1.0.48773; claude-code_2-1-246_agent)
circleci-cli (darwin/arm64; 1.0.48692; codex)
circleci-cli (darwin/arm64; 1.0.48773; opencode)
circleci mcp (755dd082f9ef571f74ac433160ab068e10c9ef25)</code></pre></div><p><span>That is self-reported identity, not inference from traffic shape. If you want to measure agent usage of your own API and you control a client, stamping the harness name into the user agent is the highest-value thing you can do. It costs one line and turns an unanswerable question into a query.</span></p><p><span>Two caveats to keep in mind if you try something similar: An agent that hand-rolls its own HTTP calls is invisible to this method. I checked the bare-HTTP-client population and found roughly 41,000 spans a day where an agent could hide indistinguishably from a shell script (the span counts come from a sampled tracing dataset: the agent-filtered queries returned unsampled, the whole-traffic ones at 1.7&#215; mean sample rate.)</span></p><h2><span>Results</span></h2><p><span>Interestingly, we realized the agents don&#8217;t care about the API directly. Instead, their focus is the MCP and CLI. Having a consistent API made it easier to build a good CLI. It&#8217;s only in this roundabout way that the agents are interested in using the API.</span></p><h3><span>Agents overwhelmingly use V3. But not by migrating.</span></h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!4Src!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ae26c2b-2ab2-47ea-b098-375b8fdabe55_1276x378.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!4Src!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ae26c2b-2ab2-47ea-b098-375b8fdabe55_1276x378.png 424w, https://substackcdn.com/image/fetch/$s_!4Src!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ae26c2b-2ab2-47ea-b098-375b8fdabe55_1276x378.png 848w, https://substackcdn.com/image/fetch/$s_!4Src!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ae26c2b-2ab2-47ea-b098-375b8fdabe55_1276x378.png 1272w, https://substackcdn.com/image/fetch/$s_!4Src!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ae26c2b-2ab2-47ea-b098-375b8fdabe55_1276x378.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!4Src!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ae26c2b-2ab2-47ea-b098-375b8fdabe55_1276x378.png" width="1276" height="378" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6ae26c2b-2ab2-47ea-b098-375b8fdabe55_1276x378.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:378,&quot;width&quot;:1276,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:50168,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/215846886?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ae26c2b-2ab2-47ea-b098-375b8fdabe55_1276x378.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!4Src!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ae26c2b-2ab2-47ea-b098-375b8fdabe55_1276x378.png 424w, https://substackcdn.com/image/fetch/$s_!4Src!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ae26c2b-2ab2-47ea-b098-375b8fdabe55_1276x378.png 848w, https://substackcdn.com/image/fetch/$s_!4Src!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ae26c2b-2ab2-47ea-b098-375b8fdabe55_1276x378.png 1272w, https://substackcdn.com/image/fetch/$s_!4Src!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6ae26c2b-2ab2-47ea-b098-375b8fdabe55_1276x378.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Hypothesis 2 confirmed, emphatically. Now the same measurement seven weeks earlier:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zi3D!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe380025c-9cb7-4940-84dd-feedd404b589_1274x374.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zi3D!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe380025c-9cb7-4940-84dd-feedd404b589_1274x374.png 424w, https://substackcdn.com/image/fetch/$s_!zi3D!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe380025c-9cb7-4940-84dd-feedd404b589_1274x374.png 848w, https://substackcdn.com/image/fetch/$s_!zi3D!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe380025c-9cb7-4940-84dd-feedd404b589_1274x374.png 1272w, https://substackcdn.com/image/fetch/$s_!zi3D!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe380025c-9cb7-4940-84dd-feedd404b589_1274x374.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zi3D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe380025c-9cb7-4940-84dd-feedd404b589_1274x374.png" width="1274" height="374" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e380025c-9cb7-4940-84dd-feedd404b589_1274x374.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:374,&quot;width&quot;:1274,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:48797,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/215846886?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe380025c-9cb7-4940-84dd-feedd404b589_1274x374.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!zi3D!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe380025c-9cb7-4940-84dd-feedd404b589_1274x374.png 424w, https://substackcdn.com/image/fetch/$s_!zi3D!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe380025c-9cb7-4940-84dd-feedd404b589_1274x374.png 848w, https://substackcdn.com/image/fetch/$s_!zi3D!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe380025c-9cb7-4940-84dd-feedd404b589_1274x374.png 1272w, https://substackcdn.com/image/fetch/$s_!zi3D!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe380025c-9cb7-4940-84dd-feedd404b589_1274x374.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>In those seven weeks, agent V3 share went from 57% to 97%. But look at the V2 column: 17,641 to 22,375. V2 agent traffic did not decline. It grew slightly. V3 did not win by converting anyone. It won by absorbing 760,000 calls of agent traffic that did not exist seven weeks earlier.</span></p><p><span>Hypothesis 3 was wrong: the migration story I expected to tell was actually a growth story.</span></p><p><span>For contrast, across all clients including browsers and scripts over that same recent week, V2 still outweighs V3 - 24,067,618 calls against 16,978,477. Humans and legacy integrations have not moved. Agents are a separate population behaving differently.</span></p><h3><span>No agents found the V3 API on its own</span></h3><p><span>Here is the entire discovery surface over thirty days. Every file we published for machines to find:</span></p><p><strong><span>Discovery endpoints &#183; All clients &#183; 30 days to 28 Aug 2026</span></strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!eB_K!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c6f1aa-7bef-4fa7-94c3-c1530d45dc88_1276x864.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!eB_K!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c6f1aa-7bef-4fa7-94c3-c1530d45dc88_1276x864.png 424w, https://substackcdn.com/image/fetch/$s_!eB_K!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c6f1aa-7bef-4fa7-94c3-c1530d45dc88_1276x864.png 848w, https://substackcdn.com/image/fetch/$s_!eB_K!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c6f1aa-7bef-4fa7-94c3-c1530d45dc88_1276x864.png 1272w, https://substackcdn.com/image/fetch/$s_!eB_K!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c6f1aa-7bef-4fa7-94c3-c1530d45dc88_1276x864.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!eB_K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c6f1aa-7bef-4fa7-94c3-c1530d45dc88_1276x864.png" width="1276" height="864" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/f1c6f1aa-7bef-4fa7-94c3-c1530d45dc88_1276x864.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:864,&quot;width&quot;:1276,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:110039,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/215846886?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c6f1aa-7bef-4fa7-94c3-c1530d45dc88_1276x864.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!eB_K!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c6f1aa-7bef-4fa7-94c3-c1530d45dc88_1276x864.png 424w, https://substackcdn.com/image/fetch/$s_!eB_K!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c6f1aa-7bef-4fa7-94c3-c1530d45dc88_1276x864.png 848w, https://substackcdn.com/image/fetch/$s_!eB_K!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c6f1aa-7bef-4fa7-94c3-c1530d45dc88_1276x864.png 1272w, https://substackcdn.com/image/fetch/$s_!eB_K!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ff1c6f1aa-7bef-4fa7-94c3-c1530d45dc88_1276x864.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Agents made 807,209 API calls in a </span><em><span>week</span></em><span>. The entire discovery surface got 6,776 hits in a </span><em><span>month</span></em><span>.</span></p><p><span>Then I broke those 6,776 hits down by user agent, expecting a long tail of coding agents. There are three populations in there, and none of them is the one I was looking for.</span></p><h3><span>Population one: nobody in particular</span></h3><p><span>Crawlers, scanners, and humans</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Amus!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94b79a3f-0df7-4a74-9baa-c3a2efa07ca2_1278x1074.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Amus!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94b79a3f-0df7-4a74-9baa-c3a2efa07ca2_1278x1074.png 424w, https://substackcdn.com/image/fetch/$s_!Amus!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94b79a3f-0df7-4a74-9baa-c3a2efa07ca2_1278x1074.png 848w, https://substackcdn.com/image/fetch/$s_!Amus!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94b79a3f-0df7-4a74-9baa-c3a2efa07ca2_1278x1074.png 1272w, https://substackcdn.com/image/fetch/$s_!Amus!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94b79a3f-0df7-4a74-9baa-c3a2efa07ca2_1278x1074.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Amus!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94b79a3f-0df7-4a74-9baa-c3a2efa07ca2_1278x1074.png" width="1278" height="1074" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/94b79a3f-0df7-4a74-9baa-c3a2efa07ca2_1278x1074.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1074,&quot;width&quot;:1278,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:162504,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/215846886?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94b79a3f-0df7-4a74-9baa-c3a2efa07ca2_1278x1074.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Amus!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94b79a3f-0df7-4a74-9baa-c3a2efa07ca2_1278x1074.png 424w, https://substackcdn.com/image/fetch/$s_!Amus!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94b79a3f-0df7-4a74-9baa-c3a2efa07ca2_1278x1074.png 848w, https://substackcdn.com/image/fetch/$s_!Amus!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94b79a3f-0df7-4a74-9baa-c3a2efa07ca2_1278x1074.png 1272w, https://substackcdn.com/image/fetch/$s_!Amus!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F94b79a3f-0df7-4a74-9baa-c3a2efa07ca2_1278x1074.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>The largest single consumer of </span><code>/.well-known/api-catalog</code><span> is a bot that exists to grade whether your API is agent-ready. We are being audited for agent-readiness by a crawler while the agents themselves never look.</span></p><h3><span>Population two: agents, but only when a human points at them</span></h3><p><span>These are visible and separable, because a directed fetch carries a different user agent than an agent doing its own work:</span></p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;49fb25e0-8245-4a5b-9f36-2c1f9e42c8a6&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">Claude-User (claude-code/2.1.220; +https://support.anthropic.com/)</code></pre></div><p><span>That is Claude Code&#8217;s web fetch, which fires when a person says &#8220;go read this URL&#8221;. Across the whole discovery surface in thirty days it accounts for 23 requests - 6 on the static OpenAPI HTML, 5 on </span><code>/docs/api/v3</code><span>, 4 on </span><code>/fullopenapi.json</code><span>, 4 on </span><code>/.well-known/api-catalog</code><span>, and single hits elsewhere. Two more came from </span><code>Claude-User/1.0</code><span>, the claude.ai equivalent.</span></p><p><span>So the honest version of the finding is not &#8220;no agent has ever read it&#8221;:</span></p><blockquote><p><strong><span>Zero agents read the discovery surface unless a human explicitly directed them to it. Twenty-three times in thirty days, someone did.</span></strong></p></blockquote><p><span>Every one of those 23 was a person (probably one of us) saying &#8220;fetch the catalog&#8221;. Not one was an agent concluding on its own that a catalog might exist and going to look.</span></p><h3><span>Population three: the training and search pipelines</span></h3><p><span>This is the part that gives me some hope. Around 203 of the 6,776 hits are corpus builders and search crawlers:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!F2TI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20655366-5637-496e-8624-3282ff75dcc4_1274x642.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!F2TI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20655366-5637-496e-8624-3282ff75dcc4_1274x642.png 424w, https://substackcdn.com/image/fetch/$s_!F2TI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20655366-5637-496e-8624-3282ff75dcc4_1274x642.png 848w, https://substackcdn.com/image/fetch/$s_!F2TI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20655366-5637-496e-8624-3282ff75dcc4_1274x642.png 1272w, https://substackcdn.com/image/fetch/$s_!F2TI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20655366-5637-496e-8624-3282ff75dcc4_1274x642.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!F2TI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20655366-5637-496e-8624-3282ff75dcc4_1274x642.png" width="1274" height="642" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/20655366-5637-496e-8624-3282ff75dcc4_1274x642.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:642,&quot;width&quot;:1274,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:77000,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/215846886?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20655366-5637-496e-8624-3282ff75dcc4_1274x642.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!F2TI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20655366-5637-496e-8624-3282ff75dcc4_1274x642.png 424w, https://substackcdn.com/image/fetch/$s_!F2TI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20655366-5637-496e-8624-3282ff75dcc4_1274x642.png 848w, https://substackcdn.com/image/fetch/$s_!F2TI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20655366-5637-496e-8624-3282ff75dcc4_1274x642.png 1272w, https://substackcdn.com/image/fetch/$s_!F2TI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F20655366-5637-496e-8624-3282ff75dcc4_1274x642.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>They are all reading </span><code>/openapi.json</code><span>. The spec </span><em><span>is</span></em><span> being ingested, but just not into a working agent&#8217;s context. It is going into the next generation of training data and the search indexes agents fall back to. The discovery surface is on a much slower clock than I assumed.</span></p><p><span>Which means the optimistic reading of this whole post is available, but it is not the one I would have guessed. It is tempting to hope that the next round of models will start consulting catalogs and </span><code>llms.txt</code><span> properly, the way the standards intend. I do not think that is what will happen. What will happen is that </span><code>GPTBot</code><span> and </span><code>ClaudeBot</code><span> read our </span><code>openapi.json</code><span>, V3 ends up in the weights, and the next generation answers V3 from memory&#8230; exactly the way this generation answers V2 from memory.</span></p><p><span>The fix for stale training data is more training data. The catalog&#8217;s delivery vehicle is the corpus, and the crawler is the courier.</span></p><p><span>So the discovery surface does pay off. Just not as discovery. And the loop it runs on is a model generation long, which is not something you can shorten by publishing harder. The only lever you control on that timescale is still the client.</span></p><h3><span>Why agents don&#8217;t discover: two layers of stale</span></h3><p><span>Watching agents work with our API, the failure has two stages and neither involves discovery.</span></p><h4><span>STAGE 1: Training data answers first</span></h4><p><span>Ask an agent about the CircleCI API and it answers immediately and confidently with V2 shapes - </span><code>/api/v2/project/{project-slug}/pipeline</code><span>, slugs in paths, the old status enum. It is not looking anything up. V2 has been in the corpus for years; V3 has existed for months.</span></p><h4><span>STAGE 2: The search index agrees with the training data</span></h4><p><span>Agents will often announce they are searching the web when they are reaching into training data. When they do genuinely search, the indexes rank a decade of V2 documentation, V2 Stack Overflow answers and V2 blog posts far above anything we published this year. Two independent mechanisms, same wrong answer.</span></p><p><span>And here is the twist that makes discovery standards nearly useless in this shape: by the time an agent finally reaches a spec or a catalog, it already believes it is looking for V2. It arrives with a target. We updated the spec descriptions to say prominently that a newer V3 API exists, and an agent that has decided it needs the V2 pipeline endpoint reads past that to go looking for the V2 pipeline endpoint.</span></p><p><span>An in-context hint only helps a reader who has not yet decided what they are looking for. An agent has always already decided.</span></p><h3><span>What actually moved the traffic</span></h3><p><span>We updated the CLI to use V3, and pointed the built-in MCP server at the same code.</span></p><p><span>My colleague Pete has written up the rebuild this came out of  (</span><a href="https://www.linkedin.com/pulse/rebuilding-circleci-cli-from-scratch-pete-steyert-woods-damje"><span>Rebuilding the CircleCI CLI from scratch</span></a><span>) and it is worth reading alongside this one, because his half of the story is the half that actually moved the numbers below.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_FhB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbb627bf-49ff-45ba-a30a-b39aff8563ba_1274x414.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_FhB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbb627bf-49ff-45ba-a30a-b39aff8563ba_1274x414.png 424w, https://substackcdn.com/image/fetch/$s_!_FhB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbb627bf-49ff-45ba-a30a-b39aff8563ba_1274x414.png 848w, https://substackcdn.com/image/fetch/$s_!_FhB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbb627bf-49ff-45ba-a30a-b39aff8563ba_1274x414.png 1272w, https://substackcdn.com/image/fetch/$s_!_FhB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbb627bf-49ff-45ba-a30a-b39aff8563ba_1274x414.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_FhB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbb627bf-49ff-45ba-a30a-b39aff8563ba_1274x414.png" width="1274" height="414" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bbb627bf-49ff-45ba-a30a-b39aff8563ba_1274x414.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:414,&quot;width&quot;:1274,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:65391,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/215846886?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbb627bf-49ff-45ba-a30a-b39aff8563ba_1274x414.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!_FhB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbb627bf-49ff-45ba-a30a-b39aff8563ba_1274x414.png 424w, https://substackcdn.com/image/fetch/$s_!_FhB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbb627bf-49ff-45ba-a30a-b39aff8563ba_1274x414.png 848w, https://substackcdn.com/image/fetch/$s_!_FhB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbb627bf-49ff-45ba-a30a-b39aff8563ba_1274x414.png 1272w, https://substackcdn.com/image/fetch/$s_!_FhB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbbb627bf-49ff-45ba-a30a-b39aff8563ba_1274x414.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Roughly 57% of all CLI traffic to the public API now carries an agent harness stamp. Every one of the 807,209 agent calls arrived through the CLI or the MCP server. None arrived from an agent constructing its own HTTP request.</span></p><p><span>Agents do not discover APIs. They use the tools already in front of them. The CLI is already installed, already authenticated, already on the agent&#8217;s </span><code>PATH</code><span>, and </span><code>--help</code><span> is right there. The MCP server appears in the tool list. Neither requires discovery, a search, or a spec.</span></p><p><strong><span>The lever was never the discovery standard. The lever was the client.</span></strong></p><h2><span>TL;DR</span></h2><p><span>No agent reads your discovery surface unless a human points at it. The catalog, </span><code>llms.txt</code><span> and the OpenAPI spec drew 6,776 hits in a month. Agent reads: 23, every one a directed fetch. Self-directed: zero.</span></p><p><span>Publish it anyway, but price it as a training-data bet. 203 of those hits were GPTBot, ClaudeBot, OAI-SearchBot, PerplexityBot and MistralAI-User. The spec is being ingested, just not into a working agent&#8217;s context.</span></p><p><span>Expect the fix to arrive through the weights, not the standard. The next model generation won&#8217;t consult your catalog. It&#8217;ll just already know your API. That&#8217;s a model-generation-long loop you can&#8217;t shorten by publishing harder.</span></p><p><span>In-context hints only help a reader who hasn&#8217;t decided yet. We put &#8220;there is a newer V3&#8221; in the spec description. An agent that has already concluded it needs a V2 endpoint reads straight past it.</span></p><p><span>Ship the client and the MCP server. 100% of measured agent traffic arrived through a wrapper that was already installed and already authenticated. None of it was hand-rolled HTTP. If you want agents on your new API, update the tools they already have.</span></p>]]></content:encoded></item><item><title><![CDATA[Your CLI docs are too long. Claude stopped reading them. ]]></title><description><![CDATA[The CircleCI CLI had good help text. Claude was only reading the first 40 lines of it.]]></description><link>https://www.confidentcommit.com/p/your-cli-docs-are-too-long-claude</link><guid isPermaLink="false">https://www.confidentcommit.com/p/your-cli-docs-are-too-long-claude</guid><dc:creator><![CDATA[Pete Steyert-Woods]]></dc:creator><pubDate>Tue, 15 Sep 2026 15:46:03 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/56211d8a-e4b4-43c9-850d-9381db9d08b9_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>I found this in a terminal transcript, not a bug report:</span></p><p><code>$ circleci run trigger --help | head -40</code></p><p><span>I hadn&#8217;t asked for </span><code>head -40</code><span>. Claude added it, and it does that constantly; help output is untrusted input of unknown size, and pulling an unbounded page into a context window is a bad trade. Forty lines </span><em><span>is</span></em><span> the interface. Below line 40 was the part that tells you how to use the command.</span></p><h2><span>Hypothesis</span></h2><p><span>As I </span><a href="https://circleci.com/blog/rebuilding-the-circleci-cli-from-scratch/"><span>shared in a previous blog post</span></a><span>, when I started rewriting the CLI, I already understood we were writing for agents; that was the premise of the rewrite. Every data-returning command has </span><code>--json</code><span> with its fields enumerated, and the command tree generates an MCP server, so each command&#8217;s description </span><em><span>is</span></em><span> a tool description.</span></p><p><span>Two things we didn&#8217;t know:</span></p><ol><li><p><strong><span>We didn&#8217;t know there was a length limit</span></strong><span>, so we never measured against one (&#8221;written for agents&#8221; was a set of content decisions, and none of them had a size).</span></p></li><li><p><strong><span>We didn&#8217;t know</span></strong><span> </span><strong><span>that we were saying everything three times</span></strong><span>: the accepted values for </span><code>--event-preset</code><span>, the default for </span><code>--provider</code><span>, all of it in the examples, in the flag table, and again in the prose. Each section had been written to stand on its own, and nobody diffs a paragraph against a table.</span></p></li></ol><p><span>So the fix wasn&#8217;t &#8220;write less&#8221;. If the length was repetition, most of it would come out mechanically; and if prose went </span><em><span>last</span></em><span>, truncation would cut the cheapest content rather than the most expensive.</span></p><h2><span>Setup</span></h2><p><span>I measured before editing anything, because &#8220;is this help text too long?&#8221; is otherwise a taste argument nobody wins. Walk the command tree, render every page the way a real invocation would, and count. Not total lines, but whether each section </span><em><span>finishes</span></em><span> inside the first 40. A flag table that starts on line 38 isn&#8217;t visible; it&#8217;s a teaser.</span></p><p><span>Across 169 command pages, the mean was 49 lines, the worst 91, and only 15% got their examples inside the window (the section an agent actually copies from).</span></p><p><span>Three passes. </span><strong><span>Boilerplate out:</span></strong><span> the banner, the per-page table re-printing the same four global flags, the repeated &#8220;Learn More&#8221; links. Seventeen lines a page, no prose rewritten. </span><strong><span>Reorder:</span></strong><span> Short &#8594; Usage &#8594; Arguments &#8594; Flags &#8594; Examples &#8594; Details, so truncation eats prose first. </span><strong><span>Trim the duplication</span></strong><span>, never dropping a flag row or an example to make space. </span><code>project trigger create</code><span> went from 91 lines to 42.</span></p><p><span>Then keep it fixed. Order and boilerplate are properties of the help template now; length is a test that asserts every page fits 40 lines, with individual caps for the twenty-five commands that genuinely can&#8217;t, and those caps only ratchet down. And since these commands are largely written by an agent working from guidelines in the repo, the rule went there too.</span></p><p><span>The PR: </span><a href="https://github.com/CircleCI-Public/circleci-cli/pull/1667"><span>https://github.com/CircleCI-Public/circleci-cli/pull/1667</span></a><span>. 449 files, about 2,600 net lines of help text deleted.</span></p><h2><span>Results</span></h2><p><span>The three passes cut both the mean and the worst case, but what really matters is whether the sections an agent actually uses (flag table and the examples) finish rendering before truncation. A page that runs to 70 lines but front-loads its flags is a better interface than a 45-line page that buries them. The reorder pass was designed around this: put the cheap content last, so if something gets cut, it isn&#8217;t the part Claude copies from.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!B8lv!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61dafbc3-23bc-417f-8b77-50ae00a86085_1270x454.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!B8lv!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61dafbc3-23bc-417f-8b77-50ae00a86085_1270x454.png 424w, https://substackcdn.com/image/fetch/$s_!B8lv!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61dafbc3-23bc-417f-8b77-50ae00a86085_1270x454.png 848w, https://substackcdn.com/image/fetch/$s_!B8lv!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61dafbc3-23bc-417f-8b77-50ae00a86085_1270x454.png 1272w, https://substackcdn.com/image/fetch/$s_!B8lv!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61dafbc3-23bc-417f-8b77-50ae00a86085_1270x454.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!B8lv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61dafbc3-23bc-417f-8b77-50ae00a86085_1270x454.png" width="1270" height="454" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/61dafbc3-23bc-417f-8b77-50ae00a86085_1270x454.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:454,&quot;width&quot;:1270,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:60506,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/215845162?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61dafbc3-23bc-417f-8b77-50ae00a86085_1270x454.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!B8lv!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61dafbc3-23bc-417f-8b77-50ae00a86085_1270x454.png 424w, https://substackcdn.com/image/fetch/$s_!B8lv!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61dafbc3-23bc-417f-8b77-50ae00a86085_1270x454.png 848w, https://substackcdn.com/image/fetch/$s_!B8lv!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61dafbc3-23bc-417f-8b77-50ae00a86085_1270x454.png 1272w, https://substackcdn.com/image/fetch/$s_!B8lv!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F61dafbc3-23bc-417f-8b77-50ae00a86085_1270x454.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>No task-success delta, because I don&#8217;t have one. What changed is the interface: every command&#8217;s complete flag list now sits inside the window an agent reads, and three in four get their examples there too, against one in seven before. The old failure mode was silent, which is why it lasted. Claude wasn&#8217;t erroring; it was succeeding at a degraded version of the task on whatever fragment of the flag list it had seen. Worse than a stack trace, because nothing tells you to go and look.</span></p><h2><span>TL;DR</span></h2><p><strong><span>&#8220;Designed for agents&#8221; is a budget, not a style.</span></strong><span> We had the content decisions right and still shipped help that couldn&#8217;t be read, because we never asked how much of it arrives.</span></p><p><strong><span>Pick one canonical place for each fact.</span></strong><span> The duplication wasn&#8217;t sloppiness; it was three self-sufficient sections maintained separately. The flag table is canonical now, and prose restating it is a defect.</span></p><p><strong><span>Make it a test.</span></strong><span> A style guide asking for 40 lines would have lasted a month; the next person adding a command, human or otherwise, has no reason to know the number exists.</span></p><p><span>The pages are better for humans now, which I didn&#8217;t expect; not because agents and people want the same thing, but because neither mistake was agent-specific. Nobody benefits from reading a flag&#8217;s values three times, and nobody was reading line 70 either.</span></p>]]></content:encoded></item><item><title><![CDATA[I ran 10,000+ configs through the compiler in 4 Minutes]]></title><description><![CDATA[How a fleet of 100 Chunk sidecars turned a 10-hour sequential validation job into a 4-minute answer for a customer facing a breaking change.]]></description><link>https://www.confidentcommit.com/p/10000-configs-validated-in-4-minutes</link><guid isPermaLink="false">https://www.confidentcommit.com/p/10000-configs-validated-in-4-minutes</guid><dc:creator><![CDATA[Makoto Mizukami]]></dc:creator><pubDate>Fri, 11 Sep 2026 20:55:20 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d3900428-e1a6-4dbc-85c2-6ba0e72cc1ca_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>As part of a larger initiative to build next-generation config tooling, </span><a href="https://discuss.circleci.com/t/breaking-changes-config-compilation-updates-september-21-2026/54719"><span>a breaking change was shipped to the config compiler</span></a><span>. The purpose? Stricter and more predictable validation gives us the reliable foundation needed to support more flexible config in the future. The problem? A customer had more than 10,000 config.yml files across their organization, and needed to know how many would break </span><em><span>before</span></em><span> the change went live, not after.</span></p><p><span>At roughly the same time, I&#8217;d been spending time considering how I might be able to use Chunk sidecars for related tasks outside of testing. I wondered: Do sidecars have the potential for a wider application?</span></p><p><span>I&#8217;d say that&#8217;s representative of my style as a whole: When conducting experiments, I don&#8217;t just think about the results, I try to think about new ways of conducting the experiment that can be applied to other experiments in the future to create additional efficiencies.</span></p><p><span>In the case of this specific experiment, I was fairly certain that the problem would be easily divided and conquered by using multiple computational resources. Without that approach, the experiment would be costly and cumbersome. Chunk sidecars provided the best route forward, and resulted in a method that I (and any other CircleCI customer) can replicate within future experiments too.</span></p><h2><span>Hypothesis</span></h2><p><span>Config validation is embarrassingly parallel. Each file is independent. There&#8217;s no shared state, no ordering requirement, no reason one file&#8217;s result should wait on another&#8217;s. If we could distribute validation across a fleet of Chunk sidecars, the wall-clock time should drop to roughly: (time to validate one file) / (number of sidecars running in parallel).</span></p><p><span>The biggest question was how to keep this cost effective from a time expenditure perspective, given that sidecars take about 30 seconds each to spin up. If each sidecar only validated a handful of files before its setup overhead dominated, the fleet wouldn&#8217;t be faster than sequential. It would just be more expensive.</span></p><h2><span>Setup</span></h2><p><span>Chunk sidecars are lightweight remote microVMs that run microbuilds alongside a developer&#8217;s work session. Each one is a full Linux environment: it can run the compiler, execute arbitrary code, and report results back.</span></p><blockquote><p><span>It was important to me to conduct this work in a way the customer could reproduce. Don&#8217;t just trust me &#8211; try the same validation on your end as well to prove that the pipeline won&#8217;t be broken when the change is introduced.</span></p><p><strong>Makoto Mizukami, Senior Field Engineer JAPAC, CircleCI</strong></p></blockquote><p><span>Knowing that the customer wouldn&#8217;t have access to our database (making the experiment potentially costly for them to run themselves) I made sure to use resources that would be available to them as well &#8211; the CircleCI API and CLI.</span></p><p><span>For this experiment, I partitioned the customer&#8217;s 10,000+ configs into batches and assigned each batch to a sidecar. Each sidecar:</span></p><ol><li><p><span>Spun up</span></p></li><li><p><span>Received its batch of projects</span></p></li><li><p><span>Walked through each project to fetch its config through CircleCI API and run it through </span><code>circleci config validate--next</code></p></li><li><p><span>Reported pass/fail results, and failure reasons if any</span></p></li></ol><p><span>The fleet size was 100 sidecars in parallel at maximum. Task orchestration, command dispatches, and result aggregations were all done by a single Claude Code agent session, allowing me to save tokens (unlike Claude subagents, which hungrily consume tokens). The work was instrumented with Honeycomb throughout: 32,359 spans across the full run.</span></p><h2><span>Results</span></h2><p><span>The fleet economics held up. Setup overhead for the 100 sidecars was 2 minutes in total &#8212; small enough that the parallelism paid off at this batch size.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NlrG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157df8f7-105e-4543-b08e-ef48e1853fa4_1274x508.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NlrG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157df8f7-105e-4543-b08e-ef48e1853fa4_1274x508.png 424w, https://substackcdn.com/image/fetch/$s_!NlrG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157df8f7-105e-4543-b08e-ef48e1853fa4_1274x508.png 848w, https://substackcdn.com/image/fetch/$s_!NlrG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157df8f7-105e-4543-b08e-ef48e1853fa4_1274x508.png 1272w, https://substackcdn.com/image/fetch/$s_!NlrG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157df8f7-105e-4543-b08e-ef48e1853fa4_1274x508.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NlrG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157df8f7-105e-4543-b08e-ef48e1853fa4_1274x508.png" width="1274" height="508" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/157df8f7-105e-4543-b08e-ef48e1853fa4_1274x508.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:508,&quot;width&quot;:1274,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:62973,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/215282952?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157df8f7-105e-4543-b08e-ef48e1853fa4_1274x508.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!NlrG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157df8f7-105e-4543-b08e-ef48e1853fa4_1274x508.png 424w, https://substackcdn.com/image/fetch/$s_!NlrG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157df8f7-105e-4543-b08e-ef48e1853fa4_1274x508.png 848w, https://substackcdn.com/image/fetch/$s_!NlrG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157df8f7-105e-4543-b08e-ef48e1853fa4_1274x508.png 1272w, https://substackcdn.com/image/fetch/$s_!NlrG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F157df8f7-105e-4543-b08e-ef48e1853fa4_1274x508.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>One finding worth noting: We also observed some specific patterns in validation failures. For example, certain series of strings tended to raise a validation error. By analyzing the result of the large-scale process further, we were also able to identify some &#8220;template&#8221; projects, from which most of the 120 configs derived.</span></p><p><span>The reason that the setup was reasonably easy was that Chunk sidecars provide snapshots. I could install a basic toolset in the sidecar, take a snapshot and apply the snapshot to the sidecars. Without that, this experiment would be impossible to conduct efficiently.</span></p><h2><span>TL;DR</span></h2><p><span>Embarrassingly parallel tasks don&#8217;t need clever algorithms. They need an environment that can run many agents at once. The compiler logic didn&#8217;t change. The validation logic didn&#8217;t change. What changed was the execution model: instead of one process working through a queue, a fleet of processes each worked through a slice.</span></p><p><span>Running them sequentially would have taken hours. A Chunk sidecar fleet did it in 4 minutes.</span></p><p><span>The thing to check before you do this: Make sure each agent knows exactly what its job is. Per-agent setup costs compound. If your task is trivial and your setup is heavy, the fleet won&#8217;t help. In this case, validation was fast enough and setup was cheap enough that the math worked. Instrument it first and check.</span></p><p><span>The customer got their answer before the breaking change shipped. It was an answer we could never have gotten in a timely manner without the massive parallelism Chunk sidecars offered. It proved my hypothesis that Chunk sidecars have multiple uses, and aren&#8217;t just limited to testing. As it turns out, sidecars are well implemented so that agents can easily leverage and consume the data, making them a perfect vehicle for experiments like this one.</span></p><p><span>We complain a lot about AI agent output. But we need to be thinking more about how we&#8217;re setting our agents up for success, and whether we&#8217;re providing them with sufficient power to do their job. Sidecars are one of the best ways to prepare agents to do great work.</span></p><p></p><h4></h4>]]></content:encoded></item><item><title><![CDATA[Pencils down: The cost of a green PR]]></title><description><![CDATA[Giving a coding agent a pre-run inventory of exactly what CI will check eliminated fix loops, driving green commit rates to 100% and rework cost to $0 on all predictable failures.]]></description><link>https://www.confidentcommit.com/p/cost-of-a-green-pr</link><guid isPermaLink="false">https://www.confidentcommit.com/p/cost-of-a-green-pr</guid><dc:creator><![CDATA[Ryan E. Hamilton]]></dc:creator><pubDate>Fri, 11 Sep 2026 20:36:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/d9a141e1-c2e8-4787-b2ab-8b6347a68619_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Imagine a student who must score </span><strong><span>100%</span></strong><span> on a final exam to pass a course. Not an A-. That would fail them. They can retake the exam as many times as they want, and every attempt comes back with painfully detailed feedback.</span></p><p><span>Then suppose that same student could peek at the test in advance and take notes.</span></p><p><span>One catch. The sneak peek does not include any of the last-minute trick questions that might show up on the real exam without warning. Those get written by &#8220;</span><strong><span>the world&#8221;</span></strong><span> the moment the exam starts.</span></p><p><span>Nobody on the teaching staff knows the trick questions in advance, and new ones can show up on a retry.</span></p><p><strong><span>Pencils down.</span></strong></p><p><span>That is coding agents and CI. Merge-ready green is the only passing grade. Outer-loop CI is the real final exam. A </span><a href="https://circleci.com/blog/chunk-sidecars/"><span>Chunk sidecar</span></a><span> is the timed practice run. Every red&#8594;green fix loop (diagnose, patch, push again) charges a little tuition: LLM dollars, CI credits, wall-clock.</span></p><p><span>I wanted to know what a green PR actually costs on a fully-automated AI coding agent run. Time. Tokens. Credits. The whole bill.</span></p><p><span>We moved the needle a bit on time and tokens. Fine. Not the story.</span></p><p><span>The story is rework going to </span><strong><span>zero</span></strong><span> on the stuff we can predict. Deliver the PR in </span><strong><span>one take</span></strong><span>. Ace it at </span><strong><span>100%</span></strong><span>. The only red you should tolerate is the unknowable.</span></p><h2><span>Hypothesis</span></h2><p><span>If you show the agent what CI will check </span><strong><span>before it writes product code</span></strong><span>, give it a page of notes, and make it pass a real practice run on a sidecar, preventable outer-loop failures should stop showing up. Fix loops should not be necessary. Green commits should hit 100%.</span></p><p><span>Not &#8220;prompt harder.&#8221; Setup for success. Ace it at 100%. One take.</span></p><h2><span>Setup</span></h2><p><span>In this lab, an A- is still red. Local lint that smiled while outer-loop CI frowned still fails. A green sidecar run that dies on the real pipeline still fails.</span></p><p><span>Two cadences. Same bar: </span><strong><span>every commit you push to outer-loop CI is green.</span></strong></p><ul><li><p><strong><span>Per-task-push:</span></strong><span> after each task clears practice run, push. Seven tasks, seven trips. Ace every one.</span></p></li><li><p><strong><span>Single-push:</span></strong><span> finish the stack, clear practice run, push once at the end. Ace that one.</span></p></li></ul><p><span>How often you push is a later lever. First you have to be equipped to push green.</span></p><h3><span>Without the sneak peek</span></h3><p><span>Three calibrations on the same AFK classic </span><a href="https://www.confidentcommit.com/p/what-snake-games-have-taught-us-about"><span>Snake</span></a><span> game build. Seven tasks. Hands off the keyboard.</span></p><p><span>No inventory of checks up front. No notes. Inner-loop checks existed, but they were split between the localhost and a sidecar.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ECcB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2253dc7f-6d05-4ad4-b566-344f1ef80338_1762x542.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ECcB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2253dc7f-6d05-4ad4-b566-344f1ef80338_1762x542.png 424w, https://substackcdn.com/image/fetch/$s_!ECcB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2253dc7f-6d05-4ad4-b566-344f1ef80338_1762x542.png 848w, https://substackcdn.com/image/fetch/$s_!ECcB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2253dc7f-6d05-4ad4-b566-344f1ef80338_1762x542.png 1272w, https://substackcdn.com/image/fetch/$s_!ECcB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2253dc7f-6d05-4ad4-b566-344f1ef80338_1762x542.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ECcB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2253dc7f-6d05-4ad4-b566-344f1ef80338_1762x542.png" width="1456" height="448" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2253dc7f-6d05-4ad4-b566-344f1ef80338_1762x542.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:448,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:112566,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/215272612?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2253dc7f-6d05-4ad4-b566-344f1ef80338_1762x542.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!ECcB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2253dc7f-6d05-4ad4-b566-344f1ef80338_1762x542.png 424w, https://substackcdn.com/image/fetch/$s_!ECcB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2253dc7f-6d05-4ad4-b566-344f1ef80338_1762x542.png 848w, https://substackcdn.com/image/fetch/$s_!ECcB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2253dc7f-6d05-4ad4-b566-344f1ef80338_1762x542.png 1272w, https://substackcdn.com/image/fetch/$s_!ECcB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2253dc7f-6d05-4ad4-b566-344f1ef80338_1762x542.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>They finished, but rework kept piling up. When we feed agents endless retries and mountains of feedback, we&#8217;re just repeating the same habits we once drilled into human developers, now adopted by their AI counterparts.</span></p><p><strong><span>*NOTE:</span></strong><span> We can make retries cheaper with a structured CI pipeline </span><code>--failure-report</code><span>, but that is </span><a href="https://www.confidentcommit.com/p/we-cut-25-of-tokens-fixing-ci"><span>its own write-up</span></a><span>. Spoiler: even if you cheapen input tokens on each retry, you still have to pay time and money to retry.</span></p><h3><span>Then we showed it the test</span></h3><p><span>Before task one, before a line of game code, the agent got a short inventory: which checks fire on the practice run, which jobs fire on the thick outer pipeline, and pointers to the scripts and configs that define them. Descriptive. Not an answer key.</span></p><p><span>It took notes (</span><code>preflight.md</code><span>). It could @-read those configs when it needed more.</span></p><p><span>In our codebase that inventory is a &#8220;CI validation manifest&#8221;, which basically tells the agent what will be graded, and where the grading logic lives, </span><strong><span>before</span></strong><span> it starts guessing from vibes.</span></p><p><span>Inner-loop checks moved fully onto the sidecar, and we stopped running the same exact lint and test loops on the localhost. Outer-loop CI stayed the honest final exam, including whatever the world still wanted to throw at it.</span></p><h2><span>Results</span></h2><p><span>Headline: </span><strong><span>0 CI fix loops</span></strong><span> on both sneak-peek arms.</span></p><ul><li><p><span>Per-task-push: </span><strong><span>7/7</span></strong><span> pushed commits green.</span></p></li><li><p><span>Single-push: the one end-of-run push went green on first contact.</span></p></li></ul><p><strong><span>Zero</span></strong><span> outer-loop CI failures on either arm.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!15ib!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3961a3ef-dc72-49ec-ae2e-e7a94676d2b5_1752x758.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!15ib!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3961a3ef-dc72-49ec-ae2e-e7a94676d2b5_1752x758.png 424w, https://substackcdn.com/image/fetch/$s_!15ib!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3961a3ef-dc72-49ec-ae2e-e7a94676d2b5_1752x758.png 848w, https://substackcdn.com/image/fetch/$s_!15ib!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3961a3ef-dc72-49ec-ae2e-e7a94676d2b5_1752x758.png 1272w, https://substackcdn.com/image/fetch/$s_!15ib!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3961a3ef-dc72-49ec-ae2e-e7a94676d2b5_1752x758.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!15ib!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3961a3ef-dc72-49ec-ae2e-e7a94676d2b5_1752x758.png" width="1456" height="630" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3961a3ef-dc72-49ec-ae2e-e7a94676d2b5_1752x758.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:630,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:160292,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/215272612?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3961a3ef-dc72-49ec-ae2e-e7a94676d2b5_1752x758.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!15ib!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3961a3ef-dc72-49ec-ae2e-e7a94676d2b5_1752x758.png 424w, https://substackcdn.com/image/fetch/$s_!15ib!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3961a3ef-dc72-49ec-ae2e-e7a94676d2b5_1752x758.png 848w, https://substackcdn.com/image/fetch/$s_!15ib!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3961a3ef-dc72-49ec-ae2e-e7a94676d2b5_1752x758.png 1272w, https://substackcdn.com/image/fetch/$s_!15ib!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3961a3ef-dc72-49ec-ae2e-e7a94676d2b5_1752x758.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Diagnosis tokens on the rework line: </span><strong><span>$0</span></strong><span>. That budget can fund first-pass product code generation instead of burning LLM dollars and CI credits on red -&gt; green archaeology.</span></p><p><span>Wall-clock did not collapse. Still about </span><strong><span>50 to 60 minutes</span></strong><span>. Coding time dominated. LLM dollars landed about </span><strong><span>$15 to $17</span></strong><span>, in line with earlier attempts. The key difference: </span><strong><span>rework nearly disappeared.</span></strong></p><p><span>Quiet rework is the same muscle </span><strong><span>Merge Efficiency Ratio</span></strong><span> names in CircleCI&#8217;s </span><a href="https://circleci.com/blog/five-takeaways-2026-q2-pulse/"><span>State of Software Delivery Q2 Pulse</span></a><span> report: how many validation cycles before a change is actually done. This lab is that problem inside one agent setup. Once green is likely, push cadence and where you place the checks are how you bake low MER into the loop.</span></p><h2><span>TL;DR</span></h2><p><strong><span>A coding agent can only prevent what we can predict. Say that out loud.</span></strong></p><p><span>Linters. Unit tests. Typechecks. Formatters. Lockfile rules. Contract tests you already wrote. The deterministic stuff with a known answer. If it is on the syllabus, the agent can study it, practice it, and clear it before you spend an outer-loop credit.</span></p><p><strong><span>It cannot prevent what the world invents after the peek is printed.</span></strong></p><p><span>Late-breaking CVEs mid-run. A registry serving a different tarball than the one you resolved an hour ago. A base image that rotated overnight. An org policy that fails a dependency you did not touch. A flaky third-party API. A secrets scanner lighting up on a fixture. A new advisory. A mirror outage. A quota. Runner drift. &#8220;Works on my machine&#8221; that is really the cloud&#8217;s clock, certs, or DNS.</span></p><p><span>Platform engineers, harness authors, and coding agents cannot predict this. It arrives when the CI pipeline is already running.</span></p><p><span>Design for what the agent can own. Leave outer-loop CI in place for the rest. Ideally one take. At most two, once the world gets a vote.</span></p><p><span>Build the agent setup so it can ship the application end-to-end with </span><strong><span>zero CI failures from known deterministic sources</span></strong><span>.</span></p><p><span>Inventory of checks. Notes. A practice surface that is real (sidecar, not theater). Then push.</span></p><p><span>Do not &#8220;prompt harder&#8221; and hope the fix loop works overtime.</span></p><p><strong><span>One take. Let&#8217;s go.</span></strong></p><p></p><h4></h4>]]></content:encoded></item><item><title><![CDATA[We cut 25% of tokens fixing CI. Here’s how]]></title><description><![CDATA[It used to be that an agent, when handed a failed CI run, spent many turns just figuring out what broke. Here&#8217;s how we changed that.]]></description><link>https://www.confidentcommit.com/p/we-cut-25-of-tokens-fixing-ci</link><guid isPermaLink="false">https://www.confidentcommit.com/p/we-cut-25-of-tokens-fixing-ci</guid><dc:creator><![CDATA[Joaquin Sandoval]]></dc:creator><pubDate>Mon, 31 Aug 2026 21:13:27 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/de0d1e5d-8650-49a2-9f7d-d2b9d0834641_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>When I joined CircleCI in December of 2025, Chunk was already in GA. At the time, we were testing its ability to offer PRs to fix broken builds. But we had 2 big problems: people weren&#8217;t using it, and the people who were, weren&#8217;t really merging the fixes that Chunk provided for their broken pipelines.</span></p><p><span>We wanted to understand: How effective was Chunk? Was there something we could do to improve the fixes given to the users? And perhaps most importantly, was there a way to improve the accessibility of the failure data to agents?</span></p><p><span>We tested the effectiveness in a controlled eval benchmark to measure the actual impact. The same dataset, the same model (claude-sonnet-5), against three different setups: no CLI, CLI with free use, and CLI with </span><code>--failure-report</code><span>. Here is what we found.</span></p><h2><span>The context</span></h2><p><span>Our first experiments were designed to mimic how Chunk works: giving it a file, then a generic prompt written by the team (for example &#8220;you are a coding agent, you need to solve this error&#8221; etc). Essentially at the beginning, we were trying to change the prompt to see if the fixes improved, and if giving certain instructions, steps, or validations would move the needle. The reality? The prompt didn&#8217;t seem to matter much.</span></p><p><span>Our team was constantly reading blog posts from other AI or CI/CD companies about what they were doing to solve this problem, and eventually we came across a </span><a href="https://arxiv.org/abs/2506.03691"><span>paper by Bytedance</span></a><span>. The paper shared that when cleaning log output, most of it is useless so the best thing to do is focus on the part of the log that contains the error.</span></p><p><span>We began to wonder: How do we remove the unnecessary characters and do a diff between the good and bad logs? That was the first successful experiment that gave us good results. It followed the way Chunk works. Because of that, we released the API of step output condensed. You fetch the output of a failed step and process to clean it, so you give better context to the agent.</span></p><p><span>From there, our focus eventually moved away from Chunk and toward the inner loop. We redesigned the experiment to mimic the usual developer flow.</span></p><blockquote><p><span>&#8220;I thought: &#8216;</span><em><span>If I was a dev working on a project and my pipeline failed, how would I fix it with an agent? Probably by copying the pipeline URL and giving it to an agent to fix. I saw that it had to follow a lot of steps to get to the pipeline output. The agent has context to the pull request, but it doesn&#8217;t necessarily have sufficient context.</span></em><span>&#8217;&#8221;</span></p><p><strong><span>Joaquin Sandoval Miramontes, Senior Software Engineer, CircleCI</span></strong></p></blockquote><h2><span>Hypothesis</span></h2><p><span>The hypothesis we arrived at was this: CI failure data is structured for humans reading a dashboard: status icons, log streams, nested UI. An agent working from that raw output burns context on navigation before it can start fixing. Therefore, restructuring the same data for agent consumption should reduce diagnostic turns and get the agent to fix faster.</span></p><h2><span>Setup</span></h2><p><span>I always try to do experiments in a structured fashion. I want to make sure that I have a way to repeat my own results over and over, and I&#8217;m particular about designing experiments with a hypothesis, expected results, and conclusions &#8211; I try to follow the scientific process wherever possible.</span></p><p><span>So here&#8217;s what we tried.</span></p><p><strong><span>New flag:</span></strong><span> </span><code>circleci run get [run_id] &#8211;failure-report</code></p><p><strong><span>Output format:</span></strong><span> condensed, organized [workflow &#8594; job &#8594; step], showing only what failed. Built for piping directly into an agent&#8217;s context window.</span></p><div class="highlighted_code_block" data-attrs="{&quot;language&quot;:&quot;plaintext&quot;,&quot;nodeId&quot;:&quot;90728749-9581-4595-a89b-4a48c72ddc87&quot;}" data-component-name="HighlightedCodeBlockToDOM"><pre class="shiki"><code class="language-plaintext">circleci run get [run_id] --failure-report | claude "fix the failing workflow"</code></pre></div><p><span>We compared three conditions on the same real failed pipelines from production:</span></p><ul><li><p><span>An agent given authenticated access to the CircleCI API</span></p></li><li><p><span>An agent given free use of the CircleCI CLI</span></p></li><li><p><span>An agent given access to the CircleCI CLI instructed to use --failure-report command</span></p></li></ul><h2><span>Results</span></h2><p><span>Efficiency</span></p><p><span>The flag delivers real efficiency gains:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1a8F!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c7fb8c6-22f5-4378-be5f-e017406d49a8_1294x574.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1a8F!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c7fb8c6-22f5-4378-be5f-e017406d49a8_1294x574.png 424w, https://substackcdn.com/image/fetch/$s_!1a8F!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c7fb8c6-22f5-4378-be5f-e017406d49a8_1294x574.png 848w, https://substackcdn.com/image/fetch/$s_!1a8F!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c7fb8c6-22f5-4378-be5f-e017406d49a8_1294x574.png 1272w, https://substackcdn.com/image/fetch/$s_!1a8F!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c7fb8c6-22f5-4378-be5f-e017406d49a8_1294x574.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1a8F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c7fb8c6-22f5-4378-be5f-e017406d49a8_1294x574.png" width="1294" height="574" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8c7fb8c6-22f5-4378-be5f-e017406d49a8_1294x574.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:574,&quot;width&quot;:1294,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:90540,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/213612700?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c7fb8c6-22f5-4378-be5f-e017406d49a8_1294x574.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1a8F!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c7fb8c6-22f5-4378-be5f-e017406d49a8_1294x574.png 424w, https://substackcdn.com/image/fetch/$s_!1a8F!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c7fb8c6-22f5-4378-be5f-e017406d49a8_1294x574.png 848w, https://substackcdn.com/image/fetch/$s_!1a8F!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c7fb8c6-22f5-4378-be5f-e017406d49a8_1294x574.png 1272w, https://substackcdn.com/image/fetch/$s_!1a8F!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c7fb8c6-22f5-4378-be5f-e017406d49a8_1294x574.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Agents using --failure-report spent fewer turns overall and finished faster. The CLI-only variant used the most turns and tokens due to agents navigating the CLI tool chain to resolve which job failed.</span></p><h4><span>Fix Quality</span></h4><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!JuOA!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471f5d23-c4e6-40a8-8803-ba1d970c450f_1308x244.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!JuOA!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471f5d23-c4e6-40a8-8803-ba1d970c450f_1308x244.png 424w, https://substackcdn.com/image/fetch/$s_!JuOA!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471f5d23-c4e6-40a8-8803-ba1d970c450f_1308x244.png 848w, https://substackcdn.com/image/fetch/$s_!JuOA!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471f5d23-c4e6-40a8-8803-ba1d970c450f_1308x244.png 1272w, https://substackcdn.com/image/fetch/$s_!JuOA!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471f5d23-c4e6-40a8-8803-ba1d970c450f_1308x244.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!JuOA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471f5d23-c4e6-40a8-8803-ba1d970c450f_1308x244.png" width="1308" height="244" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/471f5d23-c4e6-40a8-8803-ba1d970c450f_1308x244.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:244,&quot;width&quot;:1308,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:35061,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/213612700?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471f5d23-c4e6-40a8-8803-ba1d970c450f_1308x244.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!JuOA!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471f5d23-c4e6-40a8-8803-ba1d970c450f_1308x244.png 424w, https://substackcdn.com/image/fetch/$s_!JuOA!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471f5d23-c4e6-40a8-8803-ba1d970c450f_1308x244.png 848w, https://substackcdn.com/image/fetch/$s_!JuOA!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471f5d23-c4e6-40a8-8803-ba1d970c450f_1308x244.png 1272w, https://substackcdn.com/image/fetch/$s_!JuOA!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F471f5d23-c4e6-40a8-8803-ba1d970c450f_1308x244.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><span>Fix quality is flat across all variants. Fix rate difference is within the regular variation across agentic runs.</span></p><h2><span>TL;DR</span></h2><p><span>It used to be that an agent, when handed a failed CI run, spent many turns just figuring out what broke. The problem: the failure data was there, it just wasn&#8217;t accessible by an agent. The new command </span><code>circleci run get [run_id] &#8211;failure-report</code><span> changes that.</span></p><p><span>Previously, the agent was taking too many steps to get to the error itself, so we summarized it into one line. More context for the agent means better fixes, and spending less.</span></p><blockquote><p><span>&#8220;If you have repeated tasks for your agent, give them tools to get it done faster. Don&#8217;t make your agent do 3-4 things before it gets started.&#8221; <br></span><strong><span>Joaquin Sandoval Miramontes, Senior Software Engineer, CircleCI</span></strong></p></blockquote><p><span>The same failure data, restructured for agent consumption, changes how efficiently the agent works. The agent still needs to gather context from the source files, understand the failure and write the fix. What it no longer needs to do is spend 4-5 turns navigating a log hierarchy to find the failing step.</span></p><p><span>The practical conclusion: </span><strong><span>agents using </span></strong><code>--failure-report</code><strong><span> reach the same fixes in about &#8531; less time and &#188; fewer tokens. At scale, given thousands of CI fixes per day, that efficiency difference is real.</span></strong></p><p><em><span>Has your team run any similar experiments? What have you learned about restructuring failure data for agent consumption? We&#8217;d love to hear from you in the comments below.</span></em></p><div><hr></div><p><em><span>Data: 84 real CI failures from CircleCI. Model: claude-sonnet-5. Quality Judge Model: Claude-opus-4.6.</span></em></p><h4></h4>]]></content:encoded></item><item><title><![CDATA[I asked Grok 4.6 and Sonnet 5 to fix failed CI pipelines. CircleCI was the judge.]]></title><description><![CDATA[Grok 4.6 went 6 for 6 fixing broken CI pipelines at $0.22 a fix; Claude Sonnet 5 at max effort went 3 for 6 at $1.11 &#8212; and the expensive model's supposed terminal advantage never materialized.]]></description><link>https://www.confidentcommit.com/p/grok-4-vs-claude-sonnet-5-fix-failed-ci-pipelines</link><guid isPermaLink="false">https://www.confidentcommit.com/p/grok-4-vs-claude-sonnet-5-fix-failed-ci-pipelines</guid><dc:creator><![CDATA[Ryan E. Hamilton]]></dc:creator><pubDate>Wed, 26 Aug 2026 23:06:55 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ee1ccebf-c02c-4c54-95f5-8391ec483ca6_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Your pipeline goes red. You hand it to an agent. It reads the logs, writes you a confident paragraph about what broke, edits the config, triggers a rerun. You pay for every token of that loop whether CI comes back green or not.</span></p><p><span>That loop runs on my machine most days. I wanted a cheaper version of it, and I had a specific idea about where the savings were hiding.</span></p><p><span>So I broke six CircleCI pipelines on purpose. Then I saved one specific failed run of each, so every model would stare at the exact same wreck. Three AI coding setups, one instruction: fix this.</span></p><p><span>Grok 4.6 fixed all six. Twenty-two cents a fix.</span></p><p><span>Claude Sonnet 5, turned all the way up, fixed three. Every fix it did land cost about $1.11.</span></p><p><span>I walked in looking for the opposite result. The chatter said Grok would be the one to stumble on the terminal work.</span></p><h2><span>What I was actually hunting</span></h2><p><span>A routing rule, which is a boring thing with a boring name: a policy for which model gets which wreck. Teams already do this with humans. The YAML typo goes to whoever&#8217;s on rotation. The bash failure that&#8217;s been red since Tuesday goes to the one person who can actually read a wait loop. I wanted that same escalation in software, with a price tag on it. Hard shell problems to the expensive model, broken config and missing test results to the cheap one. Same quality of fix, smaller bill at the end of the month.</span></p><p><span>The stake is your inner loop. Every time an agent picks up a red pipeline, reads the logs, edits the config, triggers a run, and reads that too, somebody is getting billed. Running all of it on the expensive model, all week, is a line item.</span></p><p><strong><span>That split never showed up.</span></strong></p><p><span>What did show up is a way to measure this stuff without fooling myself. I connected </span><a href="https://cli.circleci.com/"><span>CircleCI&#8217;s command line tools</span></a><span> to both Claude Code and Cursor. I saved a red run, so nobody could grade a different pipeline than the one I picked. Then I refused to grade the model&#8217;s essay for a proposed fix. By &#8220;essay&#8221; I mean the model&#8217;s written diagnosis of what broke and what it would change, which is a completely different thing from whether CircleCI actually went green.</span></p><p><span>Agents are already good at the essay. </span><strong><span>The actual pipeline rerun is the score.</span></strong></p><p><span>The essay is the menu. Green CI is the meal. I came to eat.</span></p><h2><span>Hypothesis</span></h2><p><span>I expected Grok 4.6 to struggle on the bash-flavored failures and hold its own everywhere else.</span></p><p><span>That was the prediction. I went in with a bias: if any model was going to struggle on shell, it would be Grok 4.6 once it had to drive a terminal. Type a command, read what came back, adjust, type the next one. Shell breakage is exactly that, over and over, so that&#8217;s where I thought a gap would show.</span></p><p><span>If the reputation held, the plan was already written. Sonnet 5 at </span><code>effort: low</code><span> takes the bash cases, Grok takes the rest, and I pay less without shipping worse fixes.</span></p><p><span>Second question, cheaper to ask and just as useful: does paying for Sonnet 5 at </span><code>effort: high</code><span> buy enough extra passes to justify the receipt?</span></p><h2><span>Setup: six puzzles, three setups, one judge</span></h2><p><span>The puzzles came from the </span><a href="https://github.com/felixshiftellecon/CircleCI-Training-Koans"><span>CircleCI Training Koans</span></a><span>, a public set of short exercises that break a pipeline on purpose so you can practice un-breaking it. I used six on a throwaway project, two per flavor of broken.</span></p><p><span>Three flavors went in. Broken config, where the pipeline file itself is wrong. Tests that don&#8217;t report, where the tests genuinely run and pass but CircleCI never receives the results, so the job looks fine and tells you nothing. And shell problems: bash logic, wait loops, cache commands, container images, services.</span></p><p><span>Then I froze them: one saved failed run per puzzle, so &#8220;whatever failed most recently&#8221; couldn&#8217;t sneak in and change the question halfway through.</span></p><p><span>Three setups, same prompt every time.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EHuI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c01035b-504d-4957-a7a4-d7fbcdda8475_1296x582.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EHuI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c01035b-504d-4957-a7a4-d7fbcdda8475_1296x582.png 424w, https://substackcdn.com/image/fetch/$s_!EHuI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c01035b-504d-4957-a7a4-d7fbcdda8475_1296x582.png 848w, https://substackcdn.com/image/fetch/$s_!EHuI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c01035b-504d-4957-a7a4-d7fbcdda8475_1296x582.png 1272w, https://substackcdn.com/image/fetch/$s_!EHuI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c01035b-504d-4957-a7a4-d7fbcdda8475_1296x582.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EHuI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c01035b-504d-4957-a7a4-d7fbcdda8475_1296x582.png" width="1296" height="582" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2c01035b-504d-4957-a7a4-d7fbcdda8475_1296x582.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:582,&quot;width&quot;:1296,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:95956,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/212915417?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c01035b-504d-4957-a7a4-d7fbcdda8475_1296x582.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!EHuI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c01035b-504d-4957-a7a4-d7fbcdda8475_1296x582.png 424w, https://substackcdn.com/image/fetch/$s_!EHuI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c01035b-504d-4957-a7a4-d7fbcdda8475_1296x582.png 848w, https://substackcdn.com/image/fetch/$s_!EHuI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c01035b-504d-4957-a7a4-d7fbcdda8475_1296x582.png 1272w, https://substackcdn.com/image/fetch/$s_!EHuI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2c01035b-504d-4957-a7a4-d7fbcdda8475_1296x582.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Those dollar figures are sticker rates for tokens, not the cost of a run. What each trial actually spent shows up in Results.</span></p><p><strong><span>Changed:</span></strong><span> the setup. </span><strong><span>Held constant:</span></strong><span> the saved failed run, the prompt, the isolation, the spend ceiling.</span></p><p><span>Every trial started in a fresh clone. No resumed sessions. Sonnet never got to peek at what Grok had already worked out, and the two Sonnet effort settings didn&#8217;t share a session either. Each agent had to check its config change, then trigger a brand new CircleCI run on its own throwaway branch. Never </span><code>main</code><span>. Never the saved branch.</span></p><p><strong><span>Now the grade, which is the part that maps to your actual day.</span></strong><span> Nobody merges the model&#8217;s write-up. You merge green CI. So a trial passes only if the pipeline that agent triggered finished green. On the tests-that-don&#8217;t-report puzzles, a green job with an empty test results panel (CircleCI&#8217;s Test Summary) is still a fail, because tests that pass without reporting anything back aren&#8217;t a fix. Afterward I re-checked every run against CircleCI&#8217;s own record, so the numbers below come from CircleCI and not from an agent&#8217;s summary of its own brilliance.</span></p><p><span>I capped spending too. I watched the expensive setup burn money during calibration, took its 95th percentile, and set the ceiling at </span><strong><span>$1.25</span></strong><span> so a trial could go long without going stupid. Claude Code can hard-stop at that number. Cursor writes it down and keeps going.</span></p><p><span>Claude dollars are whatever Claude Code billed for the whole session, including the Haiku 4.5 helper that sometimes rides along. Grok dollars are tokens in and out times the published rate, not an invoice from Cursor.</span></p><p><span>One honest wrinkle: model and host are welded together here. Sonnet only runs on Claude Code, Grok 4.6 only on Cursor CLI. I can&#8217;t pull those apart, and I&#8217;d rather say it out loud than bury it in a footnote.</span></p><p><span>Eighteen isolated trials. Then a repeat round on the same six puzzles, then one genuinely messy red Playwright job. Three separate scoreboards, and averaging them would only make the numbers look tidier than the evidence is.</span></p><h2><span>Results</span></h2><p><span>The pilot, scored the day it ran.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!yctk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0888c543-1b5b-4871-b4b6-6a5f8d4c8db0_1278x602.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!yctk!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0888c543-1b5b-4871-b4b6-6a5f8d4c8db0_1278x602.png 424w, https://substackcdn.com/image/fetch/$s_!yctk!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0888c543-1b5b-4871-b4b6-6a5f8d4c8db0_1278x602.png 848w, https://substackcdn.com/image/fetch/$s_!yctk!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0888c543-1b5b-4871-b4b6-6a5f8d4c8db0_1278x602.png 1272w, https://substackcdn.com/image/fetch/$s_!yctk!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0888c543-1b5b-4871-b4b6-6a5f8d4c8db0_1278x602.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!yctk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0888c543-1b5b-4871-b4b6-6a5f8d4c8db0_1278x602.png" width="1278" height="602" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0888c543-1b5b-4871-b4b6-6a5f8d4c8db0_1278x602.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:602,&quot;width&quot;:1278,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:78785,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/212915417?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0888c543-1b5b-4871-b4b6-6a5f8d4c8db0_1278x602.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!yctk!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0888c543-1b5b-4871-b4b6-6a5f8d4c8db0_1278x602.png 424w, https://substackcdn.com/image/fetch/$s_!yctk!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0888c543-1b5b-4871-b4b6-6a5f8d4c8db0_1278x602.png 848w, https://substackcdn.com/image/fetch/$s_!yctk!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0888c543-1b5b-4871-b4b6-6a5f8d4c8db0_1278x602.png 1272w, https://substackcdn.com/image/fetch/$s_!yctk!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0888c543-1b5b-4871-b4b6-6a5f8d4c8db0_1278x602.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>All three greened both </span><strong><span>shell</span></strong><span> puzzles. The terminal gap I came hunting for never showed.</span></p><p><span>One puzzle was a config file so mangled that CircleCI rejected it before any job started. That&#8217;s why the original run was red: nothing ran at all. Fixing it meant repairing the file, triggering a new pipeline, and getting that new pipeline green. Only Grok&#8217;s new run went green. Sonnet at </span><code>effort: low</code><span> spent the entire $1.25 cap on that one puzzle and got cut off. Sonnet at </span><code>effort: high</code><span> failed it too.</span></p><p><span>Separately, Sonnet at high effort missed both puzzles where the tests ran but CircleCI never received the results. Zero for two.</span></p><p><span>I ran the same six puzzles a second time. Fresh start, no memory of the first try. Grok fixed all six again. Add both tries together and Grok is 12 for 12, Sonnet at low effort is 9 for 12, Sonnet at high effort is 7 for 12. Same homework twice. I wanted to know if Grok just got lucky on Monday. It didn&#8217;t.</span></p><p><span>Then the follow-up. I cloned a Playwright job that had been going red over and over, again a throwaway, not production. The bug was a locator: </span><code>getByText</code><span> matched both a heading and a button, so the test found two things where it wanted one and refused to guess.</span></p><p><span>All three models got that one pipeline green. One try each. Nobody deleted tests. Nobody touched </span><code>.circleci/config.yml</code><span>. Every one of them edited the Playwright test and switched to </span><code>getByRole(&#8221;heading&#8221;, ...)</code><span>. Inside each green job, CircleCI&#8217;s test results panel listed six Playwright tests, and all six passed.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!1H3b!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb19b493a-91bc-45bb-a333-0b8aab45af56_1284x402.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!1H3b!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb19b493a-91bc-45bb-a333-0b8aab45af56_1284x402.png 424w, https://substackcdn.com/image/fetch/$s_!1H3b!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb19b493a-91bc-45bb-a333-0b8aab45af56_1284x402.png 848w, https://substackcdn.com/image/fetch/$s_!1H3b!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb19b493a-91bc-45bb-a333-0b8aab45af56_1284x402.png 1272w, https://substackcdn.com/image/fetch/$s_!1H3b!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb19b493a-91bc-45bb-a333-0b8aab45af56_1284x402.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!1H3b!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb19b493a-91bc-45bb-a333-0b8aab45af56_1284x402.png" width="1284" height="402" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b19b493a-91bc-45bb-a333-0b8aab45af56_1284x402.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:402,&quot;width&quot;:1284,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:53625,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/212915417?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb19b493a-91bc-45bb-a333-0b8aab45af56_1284x402.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!1H3b!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb19b493a-91bc-45bb-a333-0b8aab45af56_1284x402.png 424w, https://substackcdn.com/image/fetch/$s_!1H3b!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb19b493a-91bc-45bb-a333-0b8aab45af56_1284x402.png 848w, https://substackcdn.com/image/fetch/$s_!1H3b!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb19b493a-91bc-45bb-a333-0b8aab45af56_1284x402.png 1272w, https://substackcdn.com/image/fetch/$s_!1H3b!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb19b493a-91bc-45bb-a333-0b8aab45af56_1284x402.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Grok was the cheapest and the slowest. High effort and low effort wrote the same fix. One case, three trials, which makes it an anecdote, and I&#8217;m labeling it one.</span></p><h2><span>TL;DR</span></h2><p><span>I didn&#8217;t get a routing policy. Two puzzles per flavor of broken is nowhere near enough to write a company rule about where your model budget goes. Grok held the shell cases, too. And paying for </span><code>effort: high</code><span> didn&#8217;t buy extra passes. On these puzzles it bought fewer.</span></p><p><span>What I did get is a test that refuses to lie to me. </span><strong><span>Most evals grade what the model wrote. This one waits for CircleCI to finish, then takes the verdict from the run itself.</span></strong><span> Execution-graded, meaning CircleCI&#8217;s pass or fail is the grade, not a rubric and not the model&#8217;s own confidence.</span></p><p><span>If you&#8217;re wiring agents into CI, that distinction is the entire ballgame.</span></p><p><strong><span>Cheap can win when execution is the judge.</span></strong><span> Twenty-two cents a successful fix against $1.11 isn&#8217;t a brand story. It&#8217;s a 5x token bill, on this lab set, every time an agent tries to get a red pipeline green. I&#8217;m not going to invent your volume. A team of ten, five of those loops a week, is about $45 extra to pay for high effort instead of Grok. That&#8217;s the unit. Not a finance-team emergency. Still a 5x.</span></p><p><span>I evaluate models for a living. Then I have to evaluate the eval.</span></p><p><span>Scoring a write-up is asking somebody how they feel. Scoring a CI pipeline rerun is taking their actual blood pressure.</span></p><p><a href="https://www.confidentcommit.com/p/we-let-an-ai-agent-say-i-passed-was"><span>RalphCI already showed that green on your laptop isn&#8217;t green in CI</span></a><span>. This one asks the ruder follow-up: hand an agent the logs, the </span><a href="https://cli.circleci.com/"><span>CircleCI tools</span></a><span>, and one saved red run, and can it actually fix the pipeline? On this lab set, yes. Often. The expensive knob wasn&#8217;t the reason.</span></p><p><span>I&#8217;m not saying always use Grok instead of Sonnet. Hosts behave differently, Cursor can&#8217;t enforce a dollar cap, and these were clean training puzzles plus one Playwright locator.</span></p><p><span>I&#8217;m saying this: </span><strong><span>if your agent loop can&#8217;t trigger a rerun and read the test results panel, you&#8217;re grading the menu instead of the meal.</span></strong></p><p></p><h4></h4>]]></content:encoded></item><item><title><![CDATA[Three habits of teams shipping 9x more validated code]]></title><description><![CDATA[Yesterday, we published our Q2 Pulse report, a follow-up to our annual State of Software Delivery that digs into a question engineering leaders are running into everywhere: why do some teams turn AI-generated code into shipped software while others watch feature branches pile up?]]></description><link>https://www.confidentcommit.com/p/three-habits-of-teams-shipping-9x</link><guid isPermaLink="false">https://www.confidentcommit.com/p/three-habits-of-teams-shipping-9x</guid><dc:creator><![CDATA[Confident Commit]]></dc:creator><pubDate>Tue, 18 Aug 2026 21:31:27 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!SE23!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b3e2da9-331e-4d93-a8aa-d39fa9273fc7_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SE23!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b3e2da9-331e-4d93-a8aa-d39fa9273fc7_1300x830.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SE23!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b3e2da9-331e-4d93-a8aa-d39fa9273fc7_1300x830.png 424w, https://substackcdn.com/image/fetch/$s_!SE23!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b3e2da9-331e-4d93-a8aa-d39fa9273fc7_1300x830.png 848w, https://substackcdn.com/image/fetch/$s_!SE23!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b3e2da9-331e-4d93-a8aa-d39fa9273fc7_1300x830.png 1272w, https://substackcdn.com/image/fetch/$s_!SE23!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b3e2da9-331e-4d93-a8aa-d39fa9273fc7_1300x830.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SE23!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b3e2da9-331e-4d93-a8aa-d39fa9273fc7_1300x830.png" width="1300" height="830" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1b3e2da9-331e-4d93-a8aa-d39fa9273fc7_1300x830.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:830,&quot;width&quot;:1300,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:41324,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/211774361?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b3e2da9-331e-4d93-a8aa-d39fa9273fc7_1300x830.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SE23!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b3e2da9-331e-4d93-a8aa-d39fa9273fc7_1300x830.png 424w, https://substackcdn.com/image/fetch/$s_!SE23!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b3e2da9-331e-4d93-a8aa-d39fa9273fc7_1300x830.png 848w, https://substackcdn.com/image/fetch/$s_!SE23!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b3e2da9-331e-4d93-a8aa-d39fa9273fc7_1300x830.png 1272w, https://substackcdn.com/image/fetch/$s_!SE23!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1b3e2da9-331e-4d93-a8aa-d39fa9273fc7_1300x830.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Yesterday, we published our<span> </span><strong><a href="https://circleci.com/resources/2026-state-of-software-delivery-q2-pulse/">Q2 Pulse report</a></strong>, a follow-up to our annual State of Software Delivery that digs into a question engineering leaders are running into everywhere: why do some teams turn AI-generated code into shipped software while others watch feature branches pile up?</p><p>AI has made writing code cheaper and faster than ever. Every team we look at is generating more of it than they were a year ago. But the data keeps showing the same split: code volume is up almost everywhere. Shipped changes are not.</p><p>The bottleneck is validation. A change goes out, a pipeline takes several minutes to run, something fails, and the cycle restarts before anything merges into main. Each round trip drains both engineering time and delivery budget. In our<span> </span><strong><a href="https://www.linkedin.com/pulse/you-paying-too-much-ship-circleci-ukt1e">previous issue</a></strong>, we modeled how a 50-developer team can spend close to<span> </span><strong>$900,000 a year</strong><span> </span>on avoidable validation cycles, driven by repeat CI runs and the tokens agents spend reloading context after slow feedback breaks the flow.</p><p>The most productive teams are pulling away because they ship in fewer cycles. That lifts individual and team velocity while driving down the cost of each change. In our Q2 Pulse report, we went behind the scenes with the most productive organizations on CircleCI to isolate the habits that help them break the delivery bottleneck without breaking the bank. Below, we&#8217;ll share three of them.</p><h3><strong>The teams setting the pace</strong></h3><p>The cohort we studied includes 20 anonymized organizations with the highest main-branch throughput on our platform. They span sizes, regions, and industries, with a concentration in security and developer tooling and additional representation from fintech, e-commerce, logistics, and energy.</p><p>As a group, they run an average of 2,165 production-branch workflows every day, a 72% jump from a year ago. That builds on the AI-driven volume surge we flagged in our Q1 report, when average daily workflow runs across CircleCI were up 59% year over year. The new finding is sharper: among the teams pushing hardest, increased activity is translating into dramatically higher production-branch throughput.</p><p>The gap between these teams and everyone else is widening. Across the platform, the most productive teams now ship roughly 9x the validated code of a typical team, up from 8x just a quarter ago. Their operational advantage comes down to three distinct habits.</p><h3><strong>Habit 1: Every engineer ships more finished work</strong></h3><p>The most productive teams get more production-ready work through the system per contributor. We measure this as main-branch workflows per contributor per day: a practical signal for how often validated work is reaching the production branch.</p><p>On a typical team, that comes out to about 1 main-branch workflow per contributor per day. Among high performers, it&#8217;s closer to 3. Among the 20 most productive teams, it&#8217;s around 12. That puts them at more than 10x the output of a typical team. It also shows how quickly the bar is rising: this same set of elite teams has nearly doubled per-contributor throughput in just one year.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!EpFc!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfed4bab-dfc5-4b59-aa9c-1d77b9c4224d_918x920.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!EpFc!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfed4bab-dfc5-4b59-aa9c-1d77b9c4224d_918x920.png 424w, https://substackcdn.com/image/fetch/$s_!EpFc!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfed4bab-dfc5-4b59-aa9c-1d77b9c4224d_918x920.png 848w, https://substackcdn.com/image/fetch/$s_!EpFc!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfed4bab-dfc5-4b59-aa9c-1d77b9c4224d_918x920.png 1272w, https://substackcdn.com/image/fetch/$s_!EpFc!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfed4bab-dfc5-4b59-aa9c-1d77b9c4224d_918x920.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!EpFc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfed4bab-dfc5-4b59-aa9c-1d77b9c4224d_918x920.png" width="918" height="920" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dfed4bab-dfc5-4b59-aa9c-1d77b9c4224d_918x920.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:920,&quot;width&quot;:918,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!EpFc!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfed4bab-dfc5-4b59-aa9c-1d77b9c4224d_918x920.png 424w, https://substackcdn.com/image/fetch/$s_!EpFc!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfed4bab-dfc5-4b59-aa9c-1d77b9c4224d_918x920.png 848w, https://substackcdn.com/image/fetch/$s_!EpFc!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfed4bab-dfc5-4b59-aa9c-1d77b9c4224d_918x920.png 1272w, https://substackcdn.com/image/fetch/$s_!EpFc!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdfed4bab-dfc5-4b59-aa9c-1d77b9c4224d_918x920.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>High per-contributor throughput is the natural result of a delivery system built around frequent integration. Branches stay short-lived, validation runs continuously, and changes do not sit for days waiting on review or late-stage cleanup.</p><p>Frequent integration matters even more with AI in the loop. As more code enters the system, the teams that pull ahead are the ones that can move changes quickly from generation to validation to merge.</p><h3><strong>Habit 2: Validation runs beyond the manual push</strong></h3><p>The mechanics behind high volume become clear when looking at how top teams trigger validation. Engineers are not manually pushing code every single time they need to find out whether something works.</p><p>Most teams only run CI when someone pushes code, and the median team triggers nearly 100% of its pipelines that way. The most productive teams do not:</p><ul><li><p><strong>Direct pushes:</strong><span> </span>Account for only about 68% of their triggers.</p></li></ul><ul><li><p><strong>Automated triggers:</strong><span> </span>The remaining 32% comes from scheduled runs, API calls, dependency-update checks, and post-deploy verification.</p></li></ul><p>Their CI keeps validating the codebase and deployment lifecycle even after devs have closed their laptops.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!9NpN!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c02f83f-cb7b-4a14-bf16-15fd3649b015_1824x956.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!9NpN!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c02f83f-cb7b-4a14-bf16-15fd3649b015_1824x956.png 424w, https://substackcdn.com/image/fetch/$s_!9NpN!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c02f83f-cb7b-4a14-bf16-15fd3649b015_1824x956.png 848w, https://substackcdn.com/image/fetch/$s_!9NpN!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c02f83f-cb7b-4a14-bf16-15fd3649b015_1824x956.png 1272w, https://substackcdn.com/image/fetch/$s_!9NpN!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c02f83f-cb7b-4a14-bf16-15fd3649b015_1824x956.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!9NpN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c02f83f-cb7b-4a14-bf16-15fd3649b015_1824x956.png" width="1456" height="763" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1c02f83f-cb7b-4a14-bf16-15fd3649b015_1824x956.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:763,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!9NpN!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c02f83f-cb7b-4a14-bf16-15fd3649b015_1824x956.png 424w, https://substackcdn.com/image/fetch/$s_!9NpN!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c02f83f-cb7b-4a14-bf16-15fd3649b015_1824x956.png 848w, https://substackcdn.com/image/fetch/$s_!9NpN!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c02f83f-cb7b-4a14-bf16-15fd3649b015_1824x956.png 1272w, https://substackcdn.com/image/fetch/$s_!9NpN!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1c02f83f-cb7b-4a14-bf16-15fd3649b015_1824x956.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>A broader trigger mix gives teams more places to catch problems, from dependency health to release verification. Done well, it can also keep push-driven feedback focused on the change a developer is trying to merge, instead of mixing that signal with unrelated failures that could have been caught elsewhere.</p><h3><strong>Habit 3: Fewer validation cycles per shipped change</strong></h3><p>Throughput and automation drive volume. Efficiency keeps that volume affordable.</p><p>We track efficiency with a metric called Merge Efficiency Ratio, or MER: the number of feature-branch workflow runs for every workflow that runs on main. A high MER means work takes more rounds of pre-merge validation before it&#8217;s ready.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NrQ6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26727677-fe24-43f6-b4ec-c11cfd93d730_884x383.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NrQ6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26727677-fe24-43f6-b4ec-c11cfd93d730_884x383.png 424w, https://substackcdn.com/image/fetch/$s_!NrQ6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26727677-fe24-43f6-b4ec-c11cfd93d730_884x383.png 848w, https://substackcdn.com/image/fetch/$s_!NrQ6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26727677-fe24-43f6-b4ec-c11cfd93d730_884x383.png 1272w, https://substackcdn.com/image/fetch/$s_!NrQ6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26727677-fe24-43f6-b4ec-c11cfd93d730_884x383.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NrQ6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26727677-fe24-43f6-b4ec-c11cfd93d730_884x383.png" width="884" height="383" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/26727677-fe24-43f6-b4ec-c11cfd93d730_884x383.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:383,&quot;width&quot;:884,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!NrQ6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26727677-fe24-43f6-b4ec-c11cfd93d730_884x383.png 424w, https://substackcdn.com/image/fetch/$s_!NrQ6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26727677-fe24-43f6-b4ec-c11cfd93d730_884x383.png 848w, https://substackcdn.com/image/fetch/$s_!NrQ6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26727677-fe24-43f6-b4ec-c11cfd93d730_884x383.png 1272w, https://substackcdn.com/image/fetch/$s_!NrQ6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F26727677-fe24-43f6-b4ec-c11cfd93d730_884x383.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>Every extra feature-branch cycle means more compute spent validating the same change. In an agentic workflow, it can also mean more tokens spent reloading context and generating fixes each time a slow run reports back. A team can ship a lot and still bleed money if every change takes five or six cycles to ship. The most productive teams keep cycles low, so their higher volume moves into production without a runaway bill behind it.</p><h3><strong>Less waiting, more shipping</strong></h3><p>Higher output per engineer, validation beyond the push, and fewer cycles per change all point to the same operating principle: the best teams reduce the distance between writing code, validating it, and merging it.</p><p>That is how more code turns into more shipped software. Feedback stays fast, branches stay short, and changes reach production instead of piling up. Just as important, the cost per change goes down instead of up.</p><p>The full Q2 Pulse report lays out the cohort behind the numbers, the implementation patterns behind each habit, and what it costs to operate at this level.</p><p><strong><a href="https://circleci.com/resources/2026-state-of-software-delivery-q2-pulse/">Read the Q2 Pulse report &#8594;</a></strong></p><p>The report focuses on the operating model behind high-throughput teams. Chunk is one practical way to start applying the same principle: move routine validation out of the slow CI loop and into the moment code is being written.</p><p><strong><a href="https://circleci.com/chunk-sidecars/">Chunk sidecars</a></strong><span> </span>run validation in the inner loop, right alongside your agent, so lint failures, broken tests, and syntax errors get caught in seconds while the context is still warm, instead of after a five-minute round trip through CI. In our own testing, microbuilds delivered feedback with up to<span> </span><strong>30x less compute usage</strong><span> </span>than a full pipeline run and about<span> </span><strong>5x fewer tokens</strong>.</p><p>Chunk sidecars are available now, free on every CircleCI plan. One command gets you started: brew install CircleCI-Public/circleci/chunk</p><p><strong><a href="https://circleci.com/blog/sidecar-microbuild-in-five-minutes/">Read the five-minute setup guide &#8594;</a></strong></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://www.confidentcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading Confident Commit! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p>]]></content:encoded></item><item><title><![CDATA[Are you paying too much to ship?]]></title><description><![CDATA[Earlier this year, the conversation in software was all about maximizing agent output. Tokenmaxxing. Jensen Huang telling engineers to spend $250K a year on tokens. Every benchmark was about throughput: how much code could your AI generate per hour?]]></description><link>https://www.confidentcommit.com/p/are-you-paying-too-much-to-ship</link><guid isPermaLink="false">https://www.confidentcommit.com/p/are-you-paying-too-much-to-ship</guid><dc:creator><![CDATA[Confident Commit]]></dc:creator><pubDate>Tue, 18 Aug 2026 21:30:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!K13B!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf515815-655f-4de9-95c5-e940328e91b2_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!K13B!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf515815-655f-4de9-95c5-e940328e91b2_1300x830.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!K13B!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf515815-655f-4de9-95c5-e940328e91b2_1300x830.png 424w, https://substackcdn.com/image/fetch/$s_!K13B!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf515815-655f-4de9-95c5-e940328e91b2_1300x830.png 848w, https://substackcdn.com/image/fetch/$s_!K13B!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf515815-655f-4de9-95c5-e940328e91b2_1300x830.png 1272w, https://substackcdn.com/image/fetch/$s_!K13B!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf515815-655f-4de9-95c5-e940328e91b2_1300x830.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!K13B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf515815-655f-4de9-95c5-e940328e91b2_1300x830.png" width="1300" height="830" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/df515815-655f-4de9-95c5-e940328e91b2_1300x830.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:830,&quot;width&quot;:1300,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:33977,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/211774248?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf515815-655f-4de9-95c5-e940328e91b2_1300x830.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!K13B!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf515815-655f-4de9-95c5-e940328e91b2_1300x830.png 424w, https://substackcdn.com/image/fetch/$s_!K13B!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf515815-655f-4de9-95c5-e940328e91b2_1300x830.png 848w, https://substackcdn.com/image/fetch/$s_!K13B!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf515815-655f-4de9-95c5-e940328e91b2_1300x830.png 1272w, https://substackcdn.com/image/fetch/$s_!K13B!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdf515815-655f-4de9-95c5-e940328e91b2_1300x830.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>Earlier this year, the conversation in software was all about maximizing agent output.<span> </span><strong><a href="https://www.forbes.com/sites/timkeary/2026/04/13/is-the-cult-of-tokenmaxxingjust-another-fad-or-the-new-normal/">Tokenmaxxing</a></strong>. Jensen Huang telling engineers to<span> </span><strong><a href="https://www.cnbc.com/2026/03/20/nvidia-ai-agents-tokens-human-workers-engineer-jobs-unemployment-jensen-huang.html">spend $250K a year on tokens</a></strong>. Every benchmark was about throughput: how much code could your AI generate per hour?</p><p>That mood has shifted fast.</p><p><strong><a href="https://fortune.com/2026/05/26/uber-coo-ai-spending-tokens-claude-code/">Uber reportedly used its entire 2026 AI coding tools budget in four months</a></strong>.<span> </span><strong><a href="https://fortune.com/2026/05/22/microsoft-ai-cost-problem-tokens-agents/">Microsoft started pulling back internal Claude Code licenses</a></strong>.<span> </span><strong><a href="https://www.businessinsider.com/amazon-ai-leaderboard-tokenmaxxing-2026-5">Amazon just scrapped an internal AI leaderboard</a></strong><span> </span>after employees started running up token costs to climb the rankings instead of shipping useful work. Across the industry, orgs are staring at massive AI bills and wondering where the ROI went.</p><p>In our<span> </span><strong><a href="https://circleci.com/blog/five-takeaways-2026-software-delivery-report/">recent research</a></strong>, we found a big part of the answer: teams are generating changes faster than ever but struggling to validate and ship them at the same pace. The delta between code written and code in production has never been wider, and it carries a hefty price tag.</p><p>In this issue, we&#8217;re digging into why teams are paying more than they should to move code into production. We&#8217;ll preview some new findings from an upcoming update to our State of Software Delivery report, show how moving validation into the inner loop of local, agent-driven development can dramatically reduce delivery time and cost, and give you practical steps to get started.</p><div><hr></div><h3><strong>Merge efficiency (and why it matters)</strong></h3><p>One of the most interesting findings from the recent State of Software Delivery was that median throughput increased 15% on feature branches but declined 7% on main branches. AI makes writing code trivial, but it adds complexity downstream when those changes hit the shared codebase.</p><p>We wanted to dig into what that burden actually looks like for a typical team. How many cycles does it take to get a piece of code into production?</p><p>To find out, we looked at the ratio of feature branch runs to main branch runs across CircleCI in March 2026. This gives us an imperfect but directionally useful proxy for the total CI overhead required to ship a single change.</p><p>Here&#8217;s what we found:</p><ul><li><p>For the median org, every main branch workflow is accompanied by<span> </span><strong>3.9 workflow runs</strong><span> </span>on a feature branch. At the mean, that number rises to<span> </span><strong>8.6</strong>. At the p95 level, it rises to<span> </span><strong>23.7</strong>.</p></li><li><p>Using main branch workflows as a rough proxy for a merged change, that means every production deploy requires<span> </span><strong>5+ total CI runs<span> </span></strong>(~4 on the feature branch and 1 on main) to complete.</p></li><li><p>Main branch workflows also fail 20-30% of the time, adding at least one more cycle when they do.</p></li></ul><p>In a human-paced development context, these numbers wouldn&#8217;t be terrible. A developer pushes a change, waits a few minutes for feedback, and iterates through a couple of build-test-fix cycles before merging to main and wrapping up the workday. That&#8217;s how you get to the canonical advice that &#8220;<strong><a href="https://martinfowler.com/articles/continuousIntegration.html#EveryonePushesCommitsToTheMainlineEveryDay">everyone pushes to main daily</a></strong>.&#8221;</p><p>In an agentic world, merge efficiency matters a lot more. When full CI runs take 5-10 minutes to deliver feedback, the agent loses context. Fixing the failure means starting a new cycle: reloading context, re-examining the change, potentially redoing work that was already completed.</p><p>That not only derails productivity but also creates massive cost inefficiencies. At 5+ CI runs for every shipped change, the math gets dire quickly.</p><div><hr></div><h3><strong>Where costs pile up</strong></h3><p>Let&#8217;s make this concrete.</p><p>Imagine an AI agent is tasked with adding a new feature to a medium-sized codebase. The full repository context, including system instructions, tool definitions, and source files, comes to about 200,000 tokens.</p><p>Modern frontier models use<span> </span><strong><a href="https://platform.claude.com/docs/en/build-with-claude/prompt-caching">prompt caching</a></strong><span> </span>to make reading massive contexts affordable. The first time the agent reads your codebase, it&#8217;s a &#8220;cold write&#8221; at a premium price. On Opus 4.7, the cost is $6.25 per million tokens. On subsequent steps, the model reads from a warm cache at a 90% discount: just $0.50 per million tokens.</p><p>But (as usual) there&#8217;s a catch: prompt caches expire after about 5 minutes of inactivity.</p><p>Here&#8217;s how that plays out in a typical delivery workflow:</p><p><strong>1. The agent loads the codebase</strong><span> </span>(200,000 tokens) and generates the code change. API cost: ~$1.28 (plus output tokens).</p><p><strong>2. The agent pushes the change</strong>, triggering a CI pipeline. The CI platform queues the job, spins up a fresh VM, pulls dependencies, compiles the project, and runs the full test suite. The pipeline consumes about 5 minutes of billable compute across its parallel jobs. But the total round trip from push to feedback, including queue time and provisioning, stretches well past the 5-minute mark.</p><p><strong>3. The prompt cache goes cold.</strong><span> </span>By the time CI reports back, the cache window has closed. The cache is wiped.</p><p><strong>4. The build fails</strong><span> </span>on a minor syntax error. The agent needs to fix it, but the cache is gone. It pays the full cold-write price to reload the codebase, and the context has grown: the previous attempt&#8217;s output and the CI build logs have pushed it from 200,000 to roughly 230,000 tokens. API cost: ~$1.44. Then the agent regenerates the fix from scratch, burning another round of output tokens at $25 per million.</p><p><strong>5. The agent pushes the fix.</strong><span> </span>Your CI provider spins up another fresh environment and runs the full suite again. Another 5 minutes of billable compute.</p><p>The merge efficiency data shows it takes an average of 5+ total CI runs to ship a single change. And with each cycle, the context grows: the agent retains the full session history, so each retry passes the original codebase plus every previous attempt along with its CI build logs. By run 5, input context can double.</p><p>In this scenario, input tokens that should cost under $2 with a warm cache balloon to over $13 across 5 cold reloads, purely because the feedback loop is slower than the agent&#8217;s cache. Add output token rework at $25 per million and $1-3 in CI compute, and the total cost per change runs to around $25.</p><p>Now imagine this at scale. A 50-developer team at agentic pace ships about 3,000 changes a month (roughly 3 per developer per working day). At ~5 CI runs per change, that&#8217;s 15,000 pipeline runs and nearly $1 million per year in token and compute costs. At 500K tokens (common for larger production codebases), it crosses $1.5 million. And with the continued increases in change volume that our data shows, those numbers climb every quarter.</p><p>Most of that spend is going to cycles that shouldn&#8217;t exist.</p><div><hr></div><h3><strong>Validation in the inner loop</strong></h3><p>Scenarios like this play out because all of the validation is happening in the outer loop: CI pipelines, shared infrastructure, billable compute. CI is the right place to catch integration failures, security issues, and anything that requires a full environment. But when a syntax error or a failing unit test has to round-trip through a full pipeline run before your agent gets the signal, you can quickly accumulate a ton of unnecessary costs.</p><p>What if those cheap failures could be caught earlier, before they ever reach CI? The way to do that is to move validation into the inner loop, where agents iterate locally before the push.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!S4Kn!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6542decb-e916-46cb-a273-190c44349190_782x504.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!S4Kn!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6542decb-e916-46cb-a273-190c44349190_782x504.png 424w, https://substackcdn.com/image/fetch/$s_!S4Kn!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6542decb-e916-46cb-a273-190c44349190_782x504.png 848w, https://substackcdn.com/image/fetch/$s_!S4Kn!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6542decb-e916-46cb-a273-190c44349190_782x504.png 1272w, https://substackcdn.com/image/fetch/$s_!S4Kn!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6542decb-e916-46cb-a273-190c44349190_782x504.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!S4Kn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6542decb-e916-46cb-a273-190c44349190_782x504.png" width="782" height="504" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/6542decb-e916-46cb-a273-190c44349190_782x504.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:504,&quot;width&quot;:782,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!S4Kn!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6542decb-e916-46cb-a273-190c44349190_782x504.png 424w, https://substackcdn.com/image/fetch/$s_!S4Kn!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6542decb-e916-46cb-a273-190c44349190_782x504.png 848w, https://substackcdn.com/image/fetch/$s_!S4Kn!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6542decb-e916-46cb-a273-190c44349190_782x504.png 1272w, https://substackcdn.com/image/fetch/$s_!S4Kn!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F6542decb-e916-46cb-a273-190c44349190_782x504.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>In a healthy system, the inner loop handles the basic checks so the outer loop only runs on code that&#8217;s ready for deeper validation. Without a quality gate before the push, your outer loop becomes a bottleneck: flooded with changes, catching failures that range from trivial to critical, burning compute and tokens on every cycle.</p><p>The path forward is to raise the quality bar in the inner loop. Validate before the push. Catch the lint failures, the broken tests, the syntax errors while the agent still has context and the fix costs almost nothing.</p><div><hr></div><h3><strong>Introducing Chunk sidecars</strong></h3><p>Moving validation into the inner loop means having an environment that can run real checks fast enough for an agent to act on the results. To get there, we recently released<span> </span><strong><a href="https://circleci.com/blog/chunk-sidecars/">Chunk sidecars</a></strong>: sandbox environments spun up by the<span> </span><strong><a href="https://github.com/CircleCI-Public/chunk-cli">Chunk CLI</a></strong><span> </span>that mirror your CI stack.</p><p>Rather than waiting for a full pipeline run to surface a failing unit test or a syntax error, the sidecar runs the validation checks you&#8217;ve configured: linting, unit tests, build validation, or whatever your stack requires. Failure feedback comes back to the agent while it still has context. The agent fixes the issue and the hook fires again, repeating until the checks pass and the change is ready to push.</p><p>This keeps the feedback loop tight enough that agents never lose context between a failure and its fix. To measure what that&#8217;s actually worth, we took real failures from the<span> </span><strong><a href="https://app.circleci.com/pipelines/gh/CircleCI-Public/chunk-cli">chunk-cli CI pipeline</a></strong><span> </span>and ran them two ways: through CI as normal, and through Chunk sidecars using microbuilds. We measured compute time and token consumption across four pipelines with warm snapshots.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DGpp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd749947c-e1dd-4ee4-ac8f-88391fd78a8e_804x469.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DGpp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd749947c-e1dd-4ee4-ac8f-88391fd78a8e_804x469.png 424w, https://substackcdn.com/image/fetch/$s_!DGpp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd749947c-e1dd-4ee4-ac8f-88391fd78a8e_804x469.png 848w, https://substackcdn.com/image/fetch/$s_!DGpp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd749947c-e1dd-4ee4-ac8f-88391fd78a8e_804x469.png 1272w, https://substackcdn.com/image/fetch/$s_!DGpp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd749947c-e1dd-4ee4-ac8f-88391fd78a8e_804x469.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DGpp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd749947c-e1dd-4ee4-ac8f-88391fd78a8e_804x469.png" width="804" height="469" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d749947c-e1dd-4ee4-ac8f-88391fd78a8e_804x469.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:469,&quot;width&quot;:804,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!DGpp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd749947c-e1dd-4ee4-ac8f-88391fd78a8e_804x469.png 424w, https://substackcdn.com/image/fetch/$s_!DGpp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd749947c-e1dd-4ee4-ac8f-88391fd78a8e_804x469.png 848w, https://substackcdn.com/image/fetch/$s_!DGpp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd749947c-e1dd-4ee4-ac8f-88391fd78a8e_804x469.png 1272w, https://substackcdn.com/image/fetch/$s_!DGpp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd749947c-e1dd-4ee4-ac8f-88391fd78a8e_804x469.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>Here&#8217;s how those numbers affect the costs we saw in the scenario introduced earlier:</p><ul><li><p><strong>Input tokens: $13 &#8594; under $2.</strong><span> </span>The 27-second feedback time keeps the agent inside the 5-minute cache window. Five cold reloads become one cold load plus four warm cache reads.</p></li><li><p><strong>Output tokens: 5 full regenerations &#8594; targeted fixes.</strong><span> </span>The agent gets a focused failure signal instead of a full CI log dump, so it applies a small fix instead of rebuilding context and regenerating the entire change.</p></li><li><p><strong>CI compute: 5+ runs &#8594; 1&#8211;2.</strong><span> </span>Fifteen to twenty minutes of billable compute per change drops to a single pipeline run on code that has already passed its checks.</p></li></ul><p>The change that was costing ~$25 in tokens and compute drops to around $6. For the 50-person team from our model, that takes annual token and compute costs from approximately $900k to around $200k.<span> </span><strong>Total savings for a 50-person team: $700k or more.</strong></p><div><hr></div><h3><strong>Stop wasting your delivery dollars</strong></h3><p>Concerns over runaway AI spend are volume problems at their core. When you scale agentic workflows without changing the feedback architecture underneath them, costs accumulate with every additional change.</p><p>Teams who have closed this loop are already separating from the pack, and that difference is widening every quarter.</p><p>If you want more return on every token and every pipeline run, look at where your validation is actually catching failures today. If most of them are happening in CI, after a push, there&#8217;s a better way.</p><p><strong><a href="https://circleci.com/blog/chunk-sidecars/">Chunk sidecars</a></strong><span> </span>are available now on all CircleCI plans, including free. Install the<span> </span><strong><a href="https://github.com/CircleCI-Public/chunk-cli">Chunk CLI</a></strong><span> </span>with brew install CircleCI-Public/circleci/chunk, run chunk init in your project, and start a sidecar session from your agent. The setup takes minutes, and the savings start on the first build.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.confidentcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.confidentcommit.com/subscribe?"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Does VCS choice determine your delivery destiny?]]></title><description><![CDATA[You can tell a lot about a developer from the tools they use.]]></description><link>https://www.confidentcommit.com/p/does-vcs-choice-determine-your-delivery</link><guid isPermaLink="false">https://www.confidentcommit.com/p/does-vcs-choice-determine-your-delivery</guid><dc:creator><![CDATA[Confident Commit]]></dc:creator><pubDate>Tue, 18 Aug 2026 21:29:20 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!vBj1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ede5dee-3ea0-4da4-8f6c-a3b1cd3194bf_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vBj1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ede5dee-3ea0-4da4-8f6c-a3b1cd3194bf_1300x830.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vBj1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ede5dee-3ea0-4da4-8f6c-a3b1cd3194bf_1300x830.png 424w, https://substackcdn.com/image/fetch/$s_!vBj1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ede5dee-3ea0-4da4-8f6c-a3b1cd3194bf_1300x830.png 848w, https://substackcdn.com/image/fetch/$s_!vBj1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ede5dee-3ea0-4da4-8f6c-a3b1cd3194bf_1300x830.png 1272w, https://substackcdn.com/image/fetch/$s_!vBj1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ede5dee-3ea0-4da4-8f6c-a3b1cd3194bf_1300x830.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vBj1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ede5dee-3ea0-4da4-8f6c-a3b1cd3194bf_1300x830.png" width="1300" height="830" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1ede5dee-3ea0-4da4-8f6c-a3b1cd3194bf_1300x830.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:830,&quot;width&quot;:1300,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:44055,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/211774164?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ede5dee-3ea0-4da4-8f6c-a3b1cd3194bf_1300x830.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vBj1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ede5dee-3ea0-4da4-8f6c-a3b1cd3194bf_1300x830.png 424w, https://substackcdn.com/image/fetch/$s_!vBj1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ede5dee-3ea0-4da4-8f6c-a3b1cd3194bf_1300x830.png 848w, https://substackcdn.com/image/fetch/$s_!vBj1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ede5dee-3ea0-4da4-8f6c-a3b1cd3194bf_1300x830.png 1272w, https://substackcdn.com/image/fetch/$s_!vBj1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1ede5dee-3ea0-4da4-8f6c-a3b1cd3194bf_1300x830.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>You can tell a lot about a developer from the tools they use. Editor, OS, language, framework. Everyone has a side, and everyone is a little suspicious of the other side&#8217;s choices. Version control is no exception. Ask a GitHub team and a GitLab team how they ship software, and you will get two pretty different descriptions of what normal looks like.</p><p>But do these preferences (and their downstream impacts) actually show up in delivery metrics? Each VCS attracts a certain kind of team, and each platform builds features to reinforce those tendencies. If teams really do self-select based on their priorities, then those priorities should show up in the delivery data.</p><p>In this issue of Confident Commit, we break down data from our most recent<span> </span><strong><a href="https://circleci.com/resources/2026-state-of-software-delivery/">State of Software Delivery</a></strong><span> </span>by code hosting platform to find out.</p><h3><strong>The evolving source of truth</strong></h3><p><strong><a href="https://git-scm.com/">Git</a></strong><span> </span>has been the foundation of collaborative software development since Linus Torvalds created it in 2005 to manage the Linux kernel. It quickly became the industry standard, displacing older centralized systems like Subversion and CVS.</p><p>Hosting platforms arrived shortly after to enable team collaboration at scale. GitHub launched in 2008, Bitbucket the same year (acquired by Atlassian in 2010), and GitLab in 2011. Between them, these three cover the vast majority of professional engineering teams today.</p><p>Despite sharing the same underlying technology, these platforms evolved to serve different kinds of teams and workflows, making unique design choices along the way:</p><ul><li><p><strong>GitHub</strong><span> </span>is built around pull-request-driven social coding, making it the natural home for high-velocity, branch-heavy workflows.</p></li><li><p><strong>Bitbucket</strong><span> </span>is tightly coupled to Jira and the Atlassian suite, attracting teams that prioritize rigorous change management and cross-functional visibility.</p></li><li><p><strong>GitLab</strong><span> </span>serves as an integrated, governance-first platform appealing to teams that require strict, centralized control over security and compliance.</p></li></ul><p>Given those philosophies, you would expect GitHub teams to iterate fastest, while Bitbucket teams move with more deliberate enterprise ceremony. GitLab teams, often working in complex and regulated environments, should land somewhere in between.</p><p>CircleCI<span> </span><strong><a href="https://circleci.com/docs/guides/integration/version-control-system-integration-overview/">integrates with all three platforms</a></strong>, so we can test those expectations directly to see if they hold up. Let&#8217;s take a look at how the hosting platforms stack up across the four key software delivery metrics.</p><h3><strong>Duration: how long is the feedback loop?</strong></h3><p>Duration measures the time from when a pipeline starts to when it finishes and determines how long a developer waits to find out whether a change passed or failed. Teams optimizing for rapid iteration want this number as low as possible. Teams running deeper validation, with more extensive tests and approval gates, accept longer durations in exchange for higher confidence in what they ship.</p><p>Workflow duration is generally consistent across all three platforms. At the median, the three cohorts are within a few seconds of each other: GitHub at 2m 58s, GitLab at 3m 4s, and Bitbucket at 2m 58s.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!ws8x!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5e6e9d-1c45-4fa2-85c3-b3f1890ed2f1_1056x628.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!ws8x!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5e6e9d-1c45-4fa2-85c3-b3f1890ed2f1_1056x628.png 424w, https://substackcdn.com/image/fetch/$s_!ws8x!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5e6e9d-1c45-4fa2-85c3-b3f1890ed2f1_1056x628.png 848w, https://substackcdn.com/image/fetch/$s_!ws8x!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5e6e9d-1c45-4fa2-85c3-b3f1890ed2f1_1056x628.png 1272w, https://substackcdn.com/image/fetch/$s_!ws8x!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5e6e9d-1c45-4fa2-85c3-b3f1890ed2f1_1056x628.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!ws8x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5e6e9d-1c45-4fa2-85c3-b3f1890ed2f1_1056x628.png" width="1056" height="628" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/da5e6e9d-1c45-4fa2-85c3-b3f1890ed2f1_1056x628.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:628,&quot;width&quot;:1056,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!ws8x!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5e6e9d-1c45-4fa2-85c3-b3f1890ed2f1_1056x628.png 424w, https://substackcdn.com/image/fetch/$s_!ws8x!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5e6e9d-1c45-4fa2-85c3-b3f1890ed2f1_1056x628.png 848w, https://substackcdn.com/image/fetch/$s_!ws8x!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5e6e9d-1c45-4fa2-85c3-b3f1890ed2f1_1056x628.png 1272w, https://substackcdn.com/image/fetch/$s_!ws8x!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda5e6e9d-1c45-4fa2-85c3-b3f1890ed2f1_1056x628.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>GitLab teams running slightly longer workflows is consistent with the hypothesis that these teams run governance-heavy pipelines, but the overall story here is similarity rather than divergence. Even at p90, pipeline durations are closely clustered at 493, 489, and 484 minutes respectively.</p><h3><strong>Throughput: how much is moving through the pipeline?</strong></h3><p>Throughput is the number of workflows a team runs per day. It can reveal how actively a team is pushing changes, and is often used as a directional indicator of team productivity.</p><p>Looking across all branches, Bitbucket teams run 1.71 workflows per day, while GitHub teams run 2.25 and GitLab teams 2.55.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HXOt!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3003fa87-0ed9-4f6b-9830-3e3ab57d5858_1200x742.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HXOt!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3003fa87-0ed9-4f6b-9830-3e3ab57d5858_1200x742.png 424w, https://substackcdn.com/image/fetch/$s_!HXOt!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3003fa87-0ed9-4f6b-9830-3e3ab57d5858_1200x742.png 848w, https://substackcdn.com/image/fetch/$s_!HXOt!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3003fa87-0ed9-4f6b-9830-3e3ab57d5858_1200x742.png 1272w, https://substackcdn.com/image/fetch/$s_!HXOt!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3003fa87-0ed9-4f6b-9830-3e3ab57d5858_1200x742.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HXOt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3003fa87-0ed9-4f6b-9830-3e3ab57d5858_1200x742.png" width="1200" height="742" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3003fa87-0ed9-4f6b-9830-3e3ab57d5858_1200x742.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:742,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!HXOt!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3003fa87-0ed9-4f6b-9830-3e3ab57d5858_1200x742.png 424w, https://substackcdn.com/image/fetch/$s_!HXOt!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3003fa87-0ed9-4f6b-9830-3e3ab57d5858_1200x742.png 848w, https://substackcdn.com/image/fetch/$s_!HXOt!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3003fa87-0ed9-4f6b-9830-3e3ab57d5858_1200x742.png 1272w, https://substackcdn.com/image/fetch/$s_!HXOt!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3003fa87-0ed9-4f6b-9830-3e3ab57d5858_1200x742.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>But the branch-level breakdown is where the numbers get interesting. Throughput on feature branches tells a story about iteration speed: how often developers are committing, running CI, and testing changes before they are ready to merge. On the main branch, throughput is a proxy for how often code moves toward production. Given what we know about each platform&#8217;s user base, you would expect GitHub teams, as the most iteration-focused cohort, to lead on throughput.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HXwo!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc98309e6-60c0-4cc6-9039-b0513a4047f9_1350x742.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HXwo!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc98309e6-60c0-4cc6-9039-b0513a4047f9_1350x742.png 424w, https://substackcdn.com/image/fetch/$s_!HXwo!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc98309e6-60c0-4cc6-9039-b0513a4047f9_1350x742.png 848w, https://substackcdn.com/image/fetch/$s_!HXwo!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc98309e6-60c0-4cc6-9039-b0513a4047f9_1350x742.png 1272w, https://substackcdn.com/image/fetch/$s_!HXwo!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc98309e6-60c0-4cc6-9039-b0513a4047f9_1350x742.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HXwo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc98309e6-60c0-4cc6-9039-b0513a4047f9_1350x742.png" width="1350" height="742" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c98309e6-60c0-4cc6-9039-b0513a4047f9_1350x742.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:742,&quot;width&quot;:1350,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!HXwo!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc98309e6-60c0-4cc6-9039-b0513a4047f9_1350x742.png 424w, https://substackcdn.com/image/fetch/$s_!HXwo!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc98309e6-60c0-4cc6-9039-b0513a4047f9_1350x742.png 848w, https://substackcdn.com/image/fetch/$s_!HXwo!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc98309e6-60c0-4cc6-9039-b0513a4047f9_1350x742.png 1272w, https://substackcdn.com/image/fetch/$s_!HXwo!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc98309e6-60c0-4cc6-9039-b0513a4047f9_1350x742.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>As expected, GitHub is the runaway leader on feature branches at 3.39 workflows per day versus 2.52 for GitLab and 1.71 for BitBucket. This makes sense considering GitHub&#8217;s user base skews toward smaller teams, open-source projects, and startups that prioritize iteration speed over formal approval processes.</p><p>But on the main branch, GitLab teams run the most workflows at 2 per day compared to 1.68 for GitHub and 1 for BitBucket. While this runs contrary to what we might first expect, the finding is consistent with<span> </span><strong><a href="https://www.linkedin.com/pulse/teams-automate-compliance-turning-regulation-hoops-competitive-gtyee/">previous research we&#8217;ve done</a></strong><span> </span>showing compliance-heavy teams shipping more frequently but hitting friction in other areas of the SDLC (more on that below).</p><h3><strong>Success rate: how healthy is the path to production?</strong></h3><p>Success rate is the percentage of workflow runs that complete successfully. The default branch is where this metric matters most. Teams with efficient delivery processes catch errors locally or on feature branches before merging, so a high success rate on main is a signal that the upstream process is working.</p><p>Average success rate on main-branch workflows is highest for Bitbucket teams at 73.7% and falls to 70.6% for GitHub and 61.3% for GitLab.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!J4xG!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92fb931a-422f-4cd6-a392-5ddcbd51871d_1200x742.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!J4xG!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92fb931a-422f-4cd6-a392-5ddcbd51871d_1200x742.png 424w, https://substackcdn.com/image/fetch/$s_!J4xG!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92fb931a-422f-4cd6-a392-5ddcbd51871d_1200x742.png 848w, https://substackcdn.com/image/fetch/$s_!J4xG!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92fb931a-422f-4cd6-a392-5ddcbd51871d_1200x742.png 1272w, https://substackcdn.com/image/fetch/$s_!J4xG!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92fb931a-422f-4cd6-a392-5ddcbd51871d_1200x742.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!J4xG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92fb931a-422f-4cd6-a392-5ddcbd51871d_1200x742.png" width="1200" height="742" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/92fb931a-422f-4cd6-a392-5ddcbd51871d_1200x742.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:742,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!J4xG!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92fb931a-422f-4cd6-a392-5ddcbd51871d_1200x742.png 424w, https://substackcdn.com/image/fetch/$s_!J4xG!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92fb931a-422f-4cd6-a392-5ddcbd51871d_1200x742.png 848w, https://substackcdn.com/image/fetch/$s_!J4xG!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92fb931a-422f-4cd6-a392-5ddcbd51871d_1200x742.png 1272w, https://substackcdn.com/image/fetch/$s_!J4xG!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F92fb931a-422f-4cd6-a392-5ddcbd51871d_1200x742.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>Bitbucket teams running the fewest pipelines but passing more of them is consistent with a stability-leaning culture where more review happens before a merge. GitHub sits in the middle, reflecting a culture comfortable operating at tempo while keeping things green. GitLab&#8217;s lower success rate likely reflects the complexity of the environments these teams operate in, something the recovery data makes even clearer.</p><h3><strong>Time to recovery: how long until you&#8217;re back to green?</strong></h3><p>Time to recovery measures resilience: how long it takes to fix a broken build and get back to a successful state.</p><p>Across all branches, GitHub teams recover in 72 minutes, Bitbucket in 80, and GitLab in 124.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!xsCI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5fbe09e-cfd4-409c-9bfb-1c003571fd19_1200x742.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!xsCI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5fbe09e-cfd4-409c-9bfb-1c003571fd19_1200x742.png 424w, https://substackcdn.com/image/fetch/$s_!xsCI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5fbe09e-cfd4-409c-9bfb-1c003571fd19_1200x742.png 848w, https://substackcdn.com/image/fetch/$s_!xsCI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5fbe09e-cfd4-409c-9bfb-1c003571fd19_1200x742.png 1272w, https://substackcdn.com/image/fetch/$s_!xsCI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5fbe09e-cfd4-409c-9bfb-1c003571fd19_1200x742.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!xsCI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5fbe09e-cfd4-409c-9bfb-1c003571fd19_1200x742.png" width="1200" height="742" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d5fbe09e-cfd4-409c-9bfb-1c003571fd19_1200x742.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:742,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!xsCI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5fbe09e-cfd4-409c-9bfb-1c003571fd19_1200x742.png 424w, https://substackcdn.com/image/fetch/$s_!xsCI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5fbe09e-cfd4-409c-9bfb-1c003571fd19_1200x742.png 848w, https://substackcdn.com/image/fetch/$s_!xsCI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5fbe09e-cfd4-409c-9bfb-1c003571fd19_1200x742.png 1272w, https://substackcdn.com/image/fetch/$s_!xsCI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd5fbe09e-cfd4-409c-9bfb-1c003571fd19_1200x742.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>On the main branch specifically, those differences are magnified: GitHub teams recover in under an hour, BitBucket teams take about 80 minutes,<span> </span><strong>and GitLab teams require nearly 8 hours to get back to green.</strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!zqYY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F808cd8fd-6ecf-469b-8f18-aefafc8f3cde_1200x742.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!zqYY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F808cd8fd-6ecf-469b-8f18-aefafc8f3cde_1200x742.png 424w, https://substackcdn.com/image/fetch/$s_!zqYY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F808cd8fd-6ecf-469b-8f18-aefafc8f3cde_1200x742.png 848w, https://substackcdn.com/image/fetch/$s_!zqYY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F808cd8fd-6ecf-469b-8f18-aefafc8f3cde_1200x742.png 1272w, https://substackcdn.com/image/fetch/$s_!zqYY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F808cd8fd-6ecf-469b-8f18-aefafc8f3cde_1200x742.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!zqYY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F808cd8fd-6ecf-469b-8f18-aefafc8f3cde_1200x742.png" width="1200" height="742" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/808cd8fd-6ecf-469b-8f18-aefafc8f3cde_1200x742.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:742,&quot;width&quot;:1200,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!zqYY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F808cd8fd-6ecf-469b-8f18-aefafc8f3cde_1200x742.png 424w, https://substackcdn.com/image/fetch/$s_!zqYY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F808cd8fd-6ecf-469b-8f18-aefafc8f3cde_1200x742.png 848w, https://substackcdn.com/image/fetch/$s_!zqYY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F808cd8fd-6ecf-469b-8f18-aefafc8f3cde_1200x742.png 1272w, https://substackcdn.com/image/fetch/$s_!zqYY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F808cd8fd-6ecf-469b-8f18-aefafc8f3cde_1200x742.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>This may seem like a large delta. However, these results are consistent with what we&#8217;ve seen previously regarding high complexity, compliance-heavy workflows. When things go sideways on main in a complex environment, the need for more extensive root cause analysis, documentation, and verification before merging a fix can add significant time to the recovery process.</p><p>Of all four metrics, time to recovery is where the self-selection pattern shows up most dramatically. The teams who chose the governance-first platform are the ones paying the governance tax on recovery.</p><h3><strong>What did we learn?</strong></h3><p>While impactful, tool choices alone are never the determining factor for how a team performs. Most of that comes down to experience, engineering vision, processes, and how you use the tools at your disposal. But it&#8217;s also true that teams gravitate toward the tools that map to their way of doing things, and over time, those choices become part of a team&#8217;s identity.</p><p>Each cohort&#8217;s delivery performance aligns closely with what their platform was designed to support. Teams self-select into the tool that matches their priorities, and those same priorities show up in the data.</p><ul><li><p><strong>GitHub teams prioritize velocity</strong>: more feature branch workflows, faster iteration, quicker recovery. This tracks with GitHub&#8217;s pull-request-driven culture where branching is cheap, experimentation is encouraged, and speed matters more than gating every change.</p></li><li><p><strong>Bitbucket teams prioritize stability</strong>: fewer workflows, highest success rate on main. The Atlassian ecosystem emphasizes structured workflows and cross-functional coordination. Teams catch more issues before merge, but that deliberation costs iteration tempo.</p></li><li><p><strong>GitLab teams prioritize governance</strong>: most main-branch workflows, longest recovery times. This reflects GitLab&#8217;s governance-first design. Teams in regulated environments often ship frequently despite their additional constraints, but when things break, the overhead of meeting compliance requirements extends recovery.</p></li></ul><p>Each platform attracts teams with different priorities, and those priorities shape delivery outcomes. So while your VCS doesn&#8217;t determine your delivery destiny, it does reflect the outcomes you&#8217;re optimizing for.</p><h3><strong>Go deeper</strong></h3><p>VCS choice is one way to see how different choices and team characteristics shape delivery performance. The<span> </span><strong><a href="https://circleci.com/blog/five-takeaways-2026-software-delivery-report/">2026 State of Software Delivery</a></strong><span> </span>takes a deeper look at what drives these differences, cutting the same metrics by team size, industry, and geography to show how delivery patterns shift depending on who you are, where you are, and what you are building.</p><p><strong><a href="https://circleci.com/resources/2026-state-of-software-delivery/">Download the report</a></strong><span> </span>to see the full analysis, or explore the data and benchmark your team against similar organizations in the<span> </span><strong><a href="https://circleci.com/software-delivery-data-explorer/">interactive Software Delivery Data Explorer</a></strong>. The more you understand about how teams like yours ship, the easier it is to know what to optimize next.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.confidentcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.confidentcommit.com/subscribe?"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[Which industries are winning the AI delivery race?]]></title><description><![CDATA[The 2026 State of Software Delivery found the largest year-over-year increase in throughput ever recorded.]]></description><link>https://www.confidentcommit.com/p/which-industries-are-winning-the</link><guid isPermaLink="false">https://www.confidentcommit.com/p/which-industries-are-winning-the</guid><dc:creator><![CDATA[Confident Commit]]></dc:creator><pubDate>Tue, 18 Aug 2026 21:28:28 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!Uyc4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5773b660-adb6-43ec-860d-7204a33d065d_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Uyc4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5773b660-adb6-43ec-860d-7204a33d065d_1300x830.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Uyc4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5773b660-adb6-43ec-860d-7204a33d065d_1300x830.png 424w, https://substackcdn.com/image/fetch/$s_!Uyc4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5773b660-adb6-43ec-860d-7204a33d065d_1300x830.png 848w, https://substackcdn.com/image/fetch/$s_!Uyc4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5773b660-adb6-43ec-860d-7204a33d065d_1300x830.png 1272w, https://substackcdn.com/image/fetch/$s_!Uyc4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5773b660-adb6-43ec-860d-7204a33d065d_1300x830.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Uyc4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5773b660-adb6-43ec-860d-7204a33d065d_1300x830.png" width="1300" height="830" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5773b660-adb6-43ec-860d-7204a33d065d_1300x830.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:830,&quot;width&quot;:1300,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:42533,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/211774009?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5773b660-adb6-43ec-860d-7204a33d065d_1300x830.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Uyc4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5773b660-adb6-43ec-860d-7204a33d065d_1300x830.png 424w, https://substackcdn.com/image/fetch/$s_!Uyc4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5773b660-adb6-43ec-860d-7204a33d065d_1300x830.png 848w, https://substackcdn.com/image/fetch/$s_!Uyc4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5773b660-adb6-43ec-860d-7204a33d065d_1300x830.png 1272w, https://substackcdn.com/image/fetch/$s_!Uyc4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5773b660-adb6-43ec-860d-7204a33d065d_1300x830.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>The<span> </span><strong><a href="https://circleci.com/resources/2026-state-of-software-delivery/">2026 State of Software Delivery</a></strong><span> </span>found the largest year-over-year increase in throughput ever recorded. Workflow volume is up 59%, but on the flip side,<span> </span><strong><a href="https://circleci.com/blog/five-takeaways-2026-software-delivery-report/">main branch activity is falling</a></strong>. Teams are generating a lot more code, but most are<span> </span><strong><a href="https://circleci.com/blog/ai-delivery-bottleneck/">struggling to turn that activity into shipped software</a></strong>.</p><p>How the AI delivery bottleneck affects your team has as much to do with your specific business requirements as it does your ability to integrate AI into your workflows. Breaking down the delivery data by industry, then, can give us useful insights into different ways the problem shows up and what teams like yours are doing to solve it.</p><p>In this issue of The Confident Commit, we&#8217;re looking at three industries that stand out in this year&#8217;s data, each for unique reasons: civil engineering, utilities, and retail. We&#8217;ll break down what each one is doing well, where they&#8217;re still falling short, and what their results can tell you about closing the gap between code generation and code delivery on your own team.</p><h3><strong>Civil engineering</strong></h3><p>Civil engineering teams validate more changes per day than any other industry. The typical project triggers 9.9 workflows every day, nearly 6x the median result across the entire dataset. On feature branches, where developers build and test new work before it&#8217;s ready for production, activity is 6.8 workflows per day.</p><p><span> </span>That high volume likely reflects how civil engineering projects are structured: multiple parallel workstreams across structural, mechanical, electrical, and site systems, each being developed simultaneously. But the picture changes when you look at the main branch, where code actually ships to production. Throughput there drops to 1.8, right in line with what most teams run.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!A-cK!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17c556d1-6819-415b-b97b-249e54363bda_1488x850.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!A-cK!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17c556d1-6819-415b-b97b-249e54363bda_1488x850.png 424w, https://substackcdn.com/image/fetch/$s_!A-cK!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17c556d1-6819-415b-b97b-249e54363bda_1488x850.png 848w, https://substackcdn.com/image/fetch/$s_!A-cK!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17c556d1-6819-415b-b97b-249e54363bda_1488x850.png 1272w, https://substackcdn.com/image/fetch/$s_!A-cK!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17c556d1-6819-415b-b97b-249e54363bda_1488x850.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!A-cK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17c556d1-6819-415b-b97b-249e54363bda_1488x850.png" width="1456" height="832" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/17c556d1-6819-415b-b97b-249e54363bda_1488x850.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:832,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!A-cK!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17c556d1-6819-415b-b97b-249e54363bda_1488x850.png 424w, https://substackcdn.com/image/fetch/$s_!A-cK!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17c556d1-6819-415b-b97b-249e54363bda_1488x850.png 848w, https://substackcdn.com/image/fetch/$s_!A-cK!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17c556d1-6819-415b-b97b-249e54363bda_1488x850.png 1272w, https://substackcdn.com/image/fetch/$s_!A-cK!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F17c556d1-6819-415b-b97b-249e54363bda_1488x850.png 1456w" sizes="100vw"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>That&#8217;s a clear illustration of the AI bottleneck in action. These teams are generating massive amounts of code, but they are paying a heavy integration tax. Their complex, interconnected systems have to be validated against each other before anything can ship. That takes time, and it doesn&#8217;t scale linearly with volume.</p><p>Given the quality and compliance constraints in this vertical, a bottleneck at the integration stage makes sense. But it also means that despite having the highest overall throughput of any industry, civil engineering&#8217;s internal velocity rarely translates into faster project delivery</p><h3><strong>Utilities</strong></h3><p>So what does it look like when a team actually breaks through the production bottleneck?</p><p>Utilities teams run 7.6 workflows per day on feature branches and 7.8 on main. Both numbers are about 4x the current global median and represent a more than 2x year-over-year increase for this vertical. But what&#8217;s especially notable is that activity on main is keeping pace with feature branches. Where civil engineering teams are increasing code volume but struggling to integrate it, utilities teams are writing<span> </span><em>and</em><span> </span>shipping more code.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lX_Y!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda3a5c68-8ae5-4fb7-a3a9-9c9b899ccd17_1488x850.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lX_Y!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda3a5c68-8ae5-4fb7-a3a9-9c9b899ccd17_1488x850.png 424w, https://substackcdn.com/image/fetch/$s_!lX_Y!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda3a5c68-8ae5-4fb7-a3a9-9c9b899ccd17_1488x850.png 848w, https://substackcdn.com/image/fetch/$s_!lX_Y!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda3a5c68-8ae5-4fb7-a3a9-9c9b899ccd17_1488x850.png 1272w, https://substackcdn.com/image/fetch/$s_!lX_Y!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda3a5c68-8ae5-4fb7-a3a9-9c9b899ccd17_1488x850.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lX_Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda3a5c68-8ae5-4fb7-a3a9-9c9b899ccd17_1488x850.png" width="1456" height="832" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/da3a5c68-8ae5-4fb7-a3a9-9c9b899ccd17_1488x850.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:832,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!lX_Y!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda3a5c68-8ae5-4fb7-a3a9-9c9b899ccd17_1488x850.png 424w, https://substackcdn.com/image/fetch/$s_!lX_Y!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda3a5c68-8ae5-4fb7-a3a9-9c9b899ccd17_1488x850.png 848w, https://substackcdn.com/image/fetch/$s_!lX_Y!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda3a5c68-8ae5-4fb7-a3a9-9c9b899ccd17_1488x850.png 1272w, https://substackcdn.com/image/fetch/$s_!lX_Y!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fda3a5c68-8ae5-4fb7-a3a9-9c9b899ccd17_1488x850.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>One possible explanation is business pressure. Utilities are critical infrastructure. Grid management, outage response, and regulatory reporting can&#8217;t wait. That urgency may push teams to ship smaller, more isolated changes that integrate cleanly rather than letting work pile up on feature branches. Civil engineering projects operate on longer timelines with less immediate consequence for any individual change, which may explain why parallel workstreams pile up and create integration complexity at main.</p><p>Whatever the reason, the data shows that utilities teams are shipping code as fast as they write it. But that&#8217;s not the complete story.</p><p>Utilities teams take approximately 400 minutes to recover from a failed build on main, compared to a global median of just 59 minutes. When something fails in a high-stakes pipeline, formal investigation steps take priority over speed regardless of how fast your CI system can rerun a build. But for a critical service that people depend on, almost 7 hours of blocked deployments has real consequences.</p><h3><strong>Retail</strong></h3><p>Like utilities, retail teams depend on 24/7 availability for their customers. While downtime for a retailer may not carry the same consequences as a utility outage, the impact on revenue and customer sentiment can be massive. That pressure has pushed retailers to optimize for uptime and fast recovery.</p><p>The typical retail team recovers from a failed build on main in 20 minutes. Most non-retail teams take 60 minutes or more.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rMAj!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878342c1-3caa-48c6-940a-d3de00435c84_1488x850.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rMAj!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878342c1-3caa-48c6-940a-d3de00435c84_1488x850.png 424w, https://substackcdn.com/image/fetch/$s_!rMAj!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878342c1-3caa-48c6-940a-d3de00435c84_1488x850.png 848w, https://substackcdn.com/image/fetch/$s_!rMAj!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878342c1-3caa-48c6-940a-d3de00435c84_1488x850.png 1272w, https://substackcdn.com/image/fetch/$s_!rMAj!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878342c1-3caa-48c6-940a-d3de00435c84_1488x850.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rMAj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878342c1-3caa-48c6-940a-d3de00435c84_1488x850.png" width="1456" height="832" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/878342c1-3caa-48c6-940a-d3de00435c84_1488x850.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:832,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!rMAj!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878342c1-3caa-48c6-940a-d3de00435c84_1488x850.png 424w, https://substackcdn.com/image/fetch/$s_!rMAj!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878342c1-3caa-48c6-940a-d3de00435c84_1488x850.png 848w, https://substackcdn.com/image/fetch/$s_!rMAj!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878342c1-3caa-48c6-940a-d3de00435c84_1488x850.png 1272w, https://substackcdn.com/image/fetch/$s_!rMAj!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F878342c1-3caa-48c6-940a-d3de00435c84_1488x850.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>Not only does retail recover the fastest, but it&#8217;s also among the top performers on success rates. Retail workflows running on main succeed 85.2% of the time, placing them fifth among all industries. With these two metrics combined, retail is the strongest all-around performer on the metrics that matter most for reliability.</p><p>Banking tells a similar story. The typical banking team recovers from a failed build in 41 minutes, second only to retail, and maintains an 80.1% success rate on main. What retail and banking have in common is that the business cost of downtime makes recovery speed an organizational priority, not just an engineering goal. For these teams, MTTR is a financial metric.</p><h3><strong>What this means for your team</strong></h3><p>Each of these industries has optimized for the thing their business pressure demands and accepted a tradeoff elsewhere. Civil engineering iterates at massive volume but ships to production at an average pace. Utilities ships to production at a very high rate, but recovery takes nearly seven hours when something breaks. Retail recovers faster than any other industry because the revenue impact of downtime made it an organizational priority.</p><p>The question for your team is which of these dimensions is actually constraining you right now. If your feature branches are active but production is flat, you have a throughput-to-production problem, the same one civil engineering is paying for at scale. If your success rate is below 80%, you&#8217;re burning engineering time and tokens on failures that stack up with every change you generate. If recovery takes more than an hour, every failure blocks not just the build that broke but everything behind it in the queue.</p><p>This is what we&#8217;re building<span> </span><strong><a href="https://circleci.com/blog/what-is-autonomous-validation/">autonomous validation</a></strong><span> </span>to solve: faster feedback, deeper context across your build history, and the ability to automatically resolve common failures before they turn into 400-minute recovery stalls or integration bottlenecks that keep code from reaching customers.</p><p><strong>See where your team stands</strong></p><p>The full<span> </span><strong><a href="https://circleci.com/resources/2026-state-of-software-delivery/">2026 State of Software Delivery</a></strong><span> </span>breaks down performance by industry, region, and team size.</p><p>Want to benchmark your own metrics against these industries? The<span> </span><strong><a href="https://circleci.com/software-delivery-data-explorer/">Software Delivery Data Explorer</a></strong><span> </span>lets you compare your throughput, success rate, duration, and recovery time against teams in your specific segment.</p><p>Ready to start closing the gap between generation and delivery? CircleCI is<span> </span><strong><a href="https://circleci.com/signup">free to get started</a></strong>.</p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.confidentcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.confidentcommit.com/subscribe?"><span>Subscribe now</span></a></p><p></p>]]></content:encoded></item><item><title><![CDATA[What 28 million workflows reveal about AI coding’s biggest risk ]]></title><description><![CDATA[In our last issue, we shared a preview of data from our upcoming 2026 State of Software Delivery showing that the promised AI productivity boom isn&#8217;t all hype.]]></description><link>https://www.confidentcommit.com/p/what-28-million-workflows-reveal</link><guid isPermaLink="false">https://www.confidentcommit.com/p/what-28-million-workflows-reveal</guid><dc:creator><![CDATA[Confident Commit]]></dc:creator><pubDate>Tue, 18 Aug 2026 21:25:52 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!kpAx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76fea7-4dc4-4475-ac8e-a9682004b789_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!kpAx!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76fea7-4dc4-4475-ac8e-a9682004b789_1300x830.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!kpAx!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76fea7-4dc4-4475-ac8e-a9682004b789_1300x830.png 424w, https://substackcdn.com/image/fetch/$s_!kpAx!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76fea7-4dc4-4475-ac8e-a9682004b789_1300x830.png 848w, https://substackcdn.com/image/fetch/$s_!kpAx!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76fea7-4dc4-4475-ac8e-a9682004b789_1300x830.png 1272w, https://substackcdn.com/image/fetch/$s_!kpAx!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76fea7-4dc4-4475-ac8e-a9682004b789_1300x830.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!kpAx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76fea7-4dc4-4475-ac8e-a9682004b789_1300x830.png" width="1300" height="830" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fb76fea7-4dc4-4475-ac8e-a9682004b789_1300x830.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:830,&quot;width&quot;:1300,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:47155,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/211773323?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76fea7-4dc4-4475-ac8e-a9682004b789_1300x830.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!kpAx!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76fea7-4dc4-4475-ac8e-a9682004b789_1300x830.png 424w, https://substackcdn.com/image/fetch/$s_!kpAx!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76fea7-4dc4-4475-ac8e-a9682004b789_1300x830.png 848w, https://substackcdn.com/image/fetch/$s_!kpAx!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76fea7-4dc4-4475-ac8e-a9682004b789_1300x830.png 1272w, https://substackcdn.com/image/fetch/$s_!kpAx!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffb76fea7-4dc4-4475-ac8e-a9682004b789_1300x830.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>In our<span> </span><strong><a href="https://www.linkedin.com/pulse/ai-productivity-boom-finally-shows-up-data-why-arent-more-teams-tiqdc/">last issue</a></strong>, we shared a preview of data from our upcoming 2026 State of Software Delivery showing that the promised AI productivity boom isn&#8217;t all hype. Throughput across the CircleCI platform increased 59% year-over-year, by far the largest productivity jump we&#8217;ve ever recorded and a clear indication that AI-assisted coding is driving massive increases in change volume.</p><p>But the gains weren&#8217;t evenly distributed. The top 5% of teams nearly doubled their output, while the median team improved by just 4%.</p><p>Since then, we&#8217;ve released the<span> </span><strong><a href="https://circleci.com/resources/2026-state-of-software-delivery/">full report</a></strong>. Alongside it, we launched the<span> </span><strong><a href="https://circleci.com/software-delivery-data-explorer/">CircleCI Software Delivery Data Explorer</a></strong>, an interactive tool built on tens of millions of real CircleCI workflows that lets you benchmark your team&#8217;s throughput, duration, success rate, and MTTR against peers in your industry, region, or team size.</p><p>In this issue, we&#8217;re digging deeper into the performance divide we called out last time. What&#8217;s actually driving this separation, and what can you do about it?</p><div><hr></div><h3><strong>The branch data tells the real story</strong></h3><p>To understand what&#8217;s causing the AI performance gap, it helps to know a little about how modern software gets built.</p><p>When a developer starts on something new, they typically don&#8217;t write code directly into the main codebase. They create a separate copy called a<span> </span><strong>branch</strong>, a safe workspace where they can build and test without affecting anything in production or blocking their team from shipping other changes. When the work is done and the code passes automated checks, it gets reviewed and merged back into that main codebase, where it can be deployed to users.</p><p>Those two events mean very different things. Activity on feature branches tells you developers are working. Activity on the main branch tells you software is shipping. When you split the throughput data along those lines, the picture changes considerably.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!lZD5!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807c9da2-0397-4c68-bfa0-20c281fe5dd4_1384x1000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!lZD5!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807c9da2-0397-4c68-bfa0-20c281fe5dd4_1384x1000.png 424w, https://substackcdn.com/image/fetch/$s_!lZD5!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807c9da2-0397-4c68-bfa0-20c281fe5dd4_1384x1000.png 848w, https://substackcdn.com/image/fetch/$s_!lZD5!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807c9da2-0397-4c68-bfa0-20c281fe5dd4_1384x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!lZD5!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807c9da2-0397-4c68-bfa0-20c281fe5dd4_1384x1000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!lZD5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807c9da2-0397-4c68-bfa0-20c281fe5dd4_1384x1000.png" width="1384" height="1000" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/807c9da2-0397-4c68-bfa0-20c281fe5dd4_1384x1000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1000,&quot;width&quot;:1384,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!lZD5!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807c9da2-0397-4c68-bfa0-20c281fe5dd4_1384x1000.png 424w, https://substackcdn.com/image/fetch/$s_!lZD5!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807c9da2-0397-4c68-bfa0-20c281fe5dd4_1384x1000.png 848w, https://substackcdn.com/image/fetch/$s_!lZD5!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807c9da2-0397-4c68-bfa0-20c281fe5dd4_1384x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!lZD5!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F807c9da2-0397-4c68-bfa0-20c281fe5dd4_1384x1000.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>For the top 5% of teams, throughput increased 26% on the main branch and 85% on feature branches between 2024 and 2025. Why is the jump so much bigger on feature branches? Because AI is particularly well-suited to the work that happens there: working out complex problems quickly, spinning up prototypes, and iterating quickly until the approach is right.</p><p>But the most important story is the increase in activity on main. Top performers are generating<span> </span><em>and shipping</em><span> </span>significantly more than they were just one year ago. That&#8217;s what it looks like when AI acceleration runs all the way through the delivery pipeline.</p><p>For everyone else, the picture is different. Teams in the top 10% saw feature branch throughput rise nearly 50%, but main branch activity was essentially flat at 1%.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7vuh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faef6276c-0a69-4a45-90fc-15d150787256_1020x1000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7vuh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faef6276c-0a69-4a45-90fc-15d150787256_1020x1000.png 424w, https://substackcdn.com/image/fetch/$s_!7vuh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faef6276c-0a69-4a45-90fc-15d150787256_1020x1000.png 848w, https://substackcdn.com/image/fetch/$s_!7vuh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faef6276c-0a69-4a45-90fc-15d150787256_1020x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!7vuh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faef6276c-0a69-4a45-90fc-15d150787256_1020x1000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7vuh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faef6276c-0a69-4a45-90fc-15d150787256_1020x1000.png" width="1020" height="1000" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/aef6276c-0a69-4a45-90fc-15d150787256_1020x1000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1000,&quot;width&quot;:1020,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!7vuh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faef6276c-0a69-4a45-90fc-15d150787256_1020x1000.png 424w, https://substackcdn.com/image/fetch/$s_!7vuh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faef6276c-0a69-4a45-90fc-15d150787256_1020x1000.png 848w, https://substackcdn.com/image/fetch/$s_!7vuh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faef6276c-0a69-4a45-90fc-15d150787256_1020x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!7vuh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Faef6276c-0a69-4a45-90fc-15d150787256_1020x1000.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>And at the median, feature branch activity was up 15%, while main branch throughput actually<span> </span><em>declined<span> </span></em>by 7%.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Q7y3!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ba6937-badc-4eba-9a96-372463bb0f4a_1004x1000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Q7y3!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ba6937-badc-4eba-9a96-372463bb0f4a_1004x1000.png 424w, https://substackcdn.com/image/fetch/$s_!Q7y3!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ba6937-badc-4eba-9a96-372463bb0f4a_1004x1000.png 848w, https://substackcdn.com/image/fetch/$s_!Q7y3!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ba6937-badc-4eba-9a96-372463bb0f4a_1004x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!Q7y3!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ba6937-badc-4eba-9a96-372463bb0f4a_1004x1000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Q7y3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ba6937-badc-4eba-9a96-372463bb0f4a_1004x1000.png" width="1004" height="1000" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/42ba6937-badc-4eba-9a96-372463bb0f4a_1004x1000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1000,&quot;width&quot;:1004,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!Q7y3!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ba6937-badc-4eba-9a96-372463bb0f4a_1004x1000.png 424w, https://substackcdn.com/image/fetch/$s_!Q7y3!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ba6937-badc-4eba-9a96-372463bb0f4a_1004x1000.png 848w, https://substackcdn.com/image/fetch/$s_!Q7y3!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ba6937-badc-4eba-9a96-372463bb0f4a_1004x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!Q7y3!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F42ba6937-badc-4eba-9a96-372463bb0f4a_1004x1000.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>In other words, the typical software team got marginally better at generating changes last year,<span> </span><strong>but they got<span> </span></strong><em><strong>worse<span> </span></strong></em><strong>at getting those changes into customers&#8217; hands.</strong></p><p>For any business leader, a regression in the organization&#8217;s ability to deliver value is going to set off alarm bells. So what&#8217;s behind it? The short answer is that AI has made it significantly easier to write code, but the process of integrating that code into a production system involves a different set of challenges. And those challenges get harder, not easier, as the volume of incoming changes goes up.</p><div><hr></div><h3><strong>Why AI-generated code is harder to land</strong></h3><p>It&#8217;s easy to say &#8220;integration is the bottleneck&#8221; and leave it there. But there are several distinct places between a developer finishing a piece of code and that code reaching production where things can slow down or break, and AI is putting pressure on all of them.</p><p>The first place AI generated changes run into trouble is code review. Most teams still require proposed changes to pass a manual review before they can be merged, and AI is making that queue harder to manage on two dimensions: volume and complexity.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!IHep!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5d3a16a-365c-4c27-960d-b0867d56f03b_1263x1000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!IHep!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5d3a16a-365c-4c27-960d-b0867d56f03b_1263x1000.png 424w, https://substackcdn.com/image/fetch/$s_!IHep!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5d3a16a-365c-4c27-960d-b0867d56f03b_1263x1000.png 848w, https://substackcdn.com/image/fetch/$s_!IHep!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5d3a16a-365c-4c27-960d-b0867d56f03b_1263x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!IHep!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5d3a16a-365c-4c27-960d-b0867d56f03b_1263x1000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!IHep!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5d3a16a-365c-4c27-960d-b0867d56f03b_1263x1000.png" width="1263" height="1000" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/e5d3a16a-365c-4c27-960d-b0867d56f03b_1263x1000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1000,&quot;width&quot;:1263,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!IHep!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5d3a16a-365c-4c27-960d-b0867d56f03b_1263x1000.png 424w, https://substackcdn.com/image/fetch/$s_!IHep!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5d3a16a-365c-4c27-960d-b0867d56f03b_1263x1000.png 848w, https://substackcdn.com/image/fetch/$s_!IHep!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5d3a16a-365c-4c27-960d-b0867d56f03b_1263x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!IHep!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fe5d3a16a-365c-4c27-960d-b0867d56f03b_1263x1000.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>We&#8217;ve already shown that AI is driving big increases in code activity for teams across the board. Those changes are not only happening more frequently, they&#8217;re also larger, denser, and touch more of the codebase than a human would in the same pass. Research by Faros AI and others has shown that AI adoption increases the size of pull requests by more than 150% on average.</p><p>What this means is that senior engineers tasked with reviewing AI-generated pull requests are facing an ever-growing review queue that is more difficult than ever to process. And they&#8217;re reporting all kinds of negative effects from it, from feelings of<span> </span><strong><a href="https://www.reddit.com/r/ExperiencedDevs/comments/1kr8clp/ai_slop_prs_are_burning_me_and_my_team_out_hard/">overwork and burnout</a></strong><span> </span>to the tendency to<span> </span><strong><a href="https://itrevolution.com/articles/the-revenge-of-qa-how-ai-code-generation-is-exposing-decades-of-process-debt/">take shortcuts and skim potentially risky code</a></strong>.</p><p>While the former is a serious risk to employee wellbeing and business continuity, only the latter is measurable in delivery data. And the impact is hard to ignore. Main branch success rates fell to<span> </span><strong>70.8%</strong><span> </span>this year, the lowest in over five years and well below the 90% benchmark that indicates a healthy pipeline.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!_CeQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c2a659e-ab3e-4ee2-8613-20b7c07f0009_1223x1000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!_CeQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c2a659e-ab3e-4ee2-8613-20b7c07f0009_1223x1000.png 424w, https://substackcdn.com/image/fetch/$s_!_CeQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c2a659e-ab3e-4ee2-8613-20b7c07f0009_1223x1000.png 848w, https://substackcdn.com/image/fetch/$s_!_CeQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c2a659e-ab3e-4ee2-8613-20b7c07f0009_1223x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!_CeQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c2a659e-ab3e-4ee2-8613-20b7c07f0009_1223x1000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!_CeQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c2a659e-ab3e-4ee2-8613-20b7c07f0009_1223x1000.png" width="1223" height="1000" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8c2a659e-ab3e-4ee2-8613-20b7c07f0009_1223x1000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1000,&quot;width&quot;:1223,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!_CeQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c2a659e-ab3e-4ee2-8613-20b7c07f0009_1223x1000.png 424w, https://substackcdn.com/image/fetch/$s_!_CeQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c2a659e-ab3e-4ee2-8613-20b7c07f0009_1223x1000.png 848w, https://substackcdn.com/image/fetch/$s_!_CeQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c2a659e-ab3e-4ee2-8613-20b7c07f0009_1223x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!_CeQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c2a659e-ab3e-4ee2-8613-20b7c07f0009_1223x1000.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>Roughly 3 in 10 attempts to merge changes into production are now failing, and when the main branch breaks, nothing else can ship until it&#8217;s fixed.</p><p>That brings us to a third place AI code tends to get stuck: fixing it when it breaks. Recovery times for failed builds have been climbing since AI-assisted development went mainstream in 2022. This year, the typical team takes<span> </span><strong>72 minutes</strong><span> </span>to get back to green, a 13% increase from the previous year.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!BG_C!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd543cc4f-476f-4832-8d6b-6de376fb19e0_1247x1000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!BG_C!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd543cc4f-476f-4832-8d6b-6de376fb19e0_1247x1000.png 424w, https://substackcdn.com/image/fetch/$s_!BG_C!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd543cc4f-476f-4832-8d6b-6de376fb19e0_1247x1000.png 848w, https://substackcdn.com/image/fetch/$s_!BG_C!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd543cc4f-476f-4832-8d6b-6de376fb19e0_1247x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!BG_C!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd543cc4f-476f-4832-8d6b-6de376fb19e0_1247x1000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!BG_C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd543cc4f-476f-4832-8d6b-6de376fb19e0_1247x1000.png" width="1247" height="1000" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d543cc4f-476f-4832-8d6b-6de376fb19e0_1247x1000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1000,&quot;width&quot;:1247,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Article content&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Article content" title="Article content" srcset="https://substackcdn.com/image/fetch/$s_!BG_C!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd543cc4f-476f-4832-8d6b-6de376fb19e0_1247x1000.png 424w, https://substackcdn.com/image/fetch/$s_!BG_C!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd543cc4f-476f-4832-8d6b-6de376fb19e0_1247x1000.png 848w, https://substackcdn.com/image/fetch/$s_!BG_C!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd543cc4f-476f-4832-8d6b-6de376fb19e0_1247x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!BG_C!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd543cc4f-476f-4832-8d6b-6de376fb19e0_1247x1000.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><p>Recovery times have been getting longer as teams incorporate more AI-generated code into their workflows. The same challenges that slow down manual review &#8211; PR size and complexity &#8211; also make failures harder to diagnose. Larger changes create more places for bugs to hide. And when something breaks in code a model generated, the developer debugging it often has little intuition for why it was written the way it was or what it might have affected downstream.</p><div><hr></div><h3><strong>CI/CD has to keep up</strong></h3><p>In a<span> </span><strong><a href="https://www.linkedin.com/pulse/building-ai-speed-why-next-leap-cicd-autonomous-validation-circleci-dntzc/">recent newsletter</a></strong>, we made the case for<span> </span><strong><a href="https://circleci.com/blog/what-is-autonomous-validation/">autonomous validation</a></strong><span> </span>as the next evolution of CI/CD. The data in this report is the clearest argument yet for why that evolution is necessary.</p><p>The AI choke points we described are all symptoms of the same underlying weakness: delivery infrastructure built for a world where humans produced code at a human pace. That&#8217;s why we&#8217;re evolving CircleCI to deliver</p><ul><li><p><strong>Faster feedback</strong><span> </span>so issues surface while context is still fresh, not after developers have moved on to the next AI-generated change</p></li><li><p><strong>Deeper context</strong><span> </span>across your build history, dependencies, configs, and policies to surface risk in AI-generated changes that a diff alone can&#8217;t reveal</p></li><li><p><strong>Agentic capabilities</strong><span> </span>that can resolve common, repeatable failures autonomously, without pulling an engineer off their work every time a flaky test or dependency conflict blocks the pipeline</p></li></ul><p>When delivery pipelines can&#8217;t keep up, AI code is more likely to stall at review, break on merge, and burn engineering time on recovery. By rethinking CI/CD to operate at the same speed and scale as AI development, we&#8217;re taking concrete steps to remove the sources of friction that keep most teams from turning AI-generated code into shipped software.</p><div><hr></div><h3><strong>Where do you stand?</strong></h3><p>If your team is generating more code than ever but struggling to get it across the finish line, you&#8217;re not alone. That&#8217;s the reality for the majority of teams in this year&#8217;s data.</p><p>You can learn more in the 2026 State of Software Delivery. It breaks down exactly where things are slowing down and what high performers are doing differently.</p><p><strong><a href="https://circleci.com/resources/2026-state-of-software-delivery/">Download the report &#8594;</a></strong></p><p>Want to see where your own numbers land? The Data Explorer lets you benchmark your delivery metrics against teams in your industry, region, and size band.</p><p><strong><a href="https://circleci.com/software-delivery-data-explorer/">Benchmark your team in the Data Explorer &#8594;</a></strong></p><p>And when you&#8217;re ready to turn all that AI-generated code into shipped software, CircleCI is free to get started.</p><p><strong><a href="https://circleci.com/signup">Start building for free &#8594;</a></strong></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.confidentcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.confidentcommit.com/subscribe?"><span>Subscribe now</span></a></p>]]></content:encoded></item><item><title><![CDATA[Junior developers have one big advantage right now]]></title><description><![CDATA[Listen now (31 mins) | Rob Zuber sits down with CircleCI engineers Michael Webster and Hanabel Mengistu to talk about what AI is actually changing about the job of software development, and what it isn't.]]></description><link>https://www.confidentcommit.com/p/junior-developers-have-one-big-advantage-a5e</link><guid isPermaLink="false">https://www.confidentcommit.com/p/junior-developers-have-one-big-advantage-a5e</guid><dc:creator><![CDATA[Confident Commit]]></dc:creator><pubDate>Thu, 06 Aug 2026 13:00:00 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/211758773/1450b5bb7ddf53dd294619053e505b5c.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><span>For the past year or two, a certain take has been making the rounds in engineering circles: we don&#8217;t need junior developers anymore.</span></p><p><span>If AI can write the code, why train someone when the model can just do it?</span></p><p><span>Rob Zuber has heard this take. He thinks it&#8217;s wrong, but not in the hand-wavy &#8220;humans will always matter&#8221; sense. He thinks it&#8217;s wrong because it misidentifies what the job actually is, and therefore misidentifies what we&#8217;d be losing.</span></p><p><span>In this episode of Confident Commit, Rob sat down with two colleagues from CircleCI&#8217;s software factory team: Michael Webster, a software engineer with more than a decade in the industry, and Hanabel Mengistu, who joined CircleCI less than a year ago straight out of college. The result is one of the more honest conversations about the current state of software development: what experienced engineers are struggling to let go of, what new engineers are struggling to find, and where those two things collide.</span></p><h2><span>1. The tools changed faster than the job descriptions did</span></h2><p><span>2025 brought a lot of change, fast.</span></p><p><span>Webster framed the inflection point precisely: &#8220;Around the end of last year, we kinda had that Sonnet 4 / Opus 4 moment where it&#8217;s like, hey, something&#8217;s happened with these models. They&#8217;re able to do a lot more.&#8221;</span></p><p><span>The instinct, for a team that wanted to push hard on AI-assisted development, was to try running multiple agents in parallel and see what broke. The answer came fast.</span></p><blockquote><p><span>&#8220;I think within the first forty-five minutes to an hour of starting, a lot of our instincts immediately started breaking down. The merge conflicts, the trying to stack them, the manual review overhead. Sub two hours, we were like, yeah, we&#8217;re gonna have to figure some things out here. Code generation time outpacing CI build times, all of that was pretty eye-opening.&#8221;</span></p><p><strong><span>Michael Webster, Principal Engineer, CircleCI</span></strong></p></blockquote><p><span>The problem wasn&#8217;t the AI. The problem was that the existing process, designed around the constraint of human coding speed, was suddenly the wrong process.</span></p><p><span>Rob&#8217;s framing: &#8220;We fundamentally changed our constraints. But as humans, we&#8217;re not good at massive leaps in approach. So we&#8217;re clinging to the processes that we have and trying to tweak them with this fundamentally new toolkit instead of asking: what was it we were trying to do, exactly?&#8221;</span></p><h2><span>2. What experienced engineers are actually clinging to (and why it&#8217;s hard)</span></h2><p><span>Webster put something into words that most senior engineers are circling around but haven&#8217;t quite said directly:</span></p><blockquote><p><span>&#8220;What if the stuff that I like doing is the bottleneck? What if the stuff that has saved me production incidents and helped me build a career, what if that is now the bottleneck? That&#8217;s a really challenging thing to work through.&#8221;</span></p><p><strong><span>Michael Webster, Principal Engineer, CircleCI</span></strong></p></blockquote><p><span>If your expertise is in a thing that&#8217;s now being commoditized, the rational move (reprioritize, go further up the stack) is emotionally costly in a way that&#8217;s hard to explain to someone who hasn&#8217;t been through it.</span></p><p><span>Webster&#8217;s working theory: go back further than you think you need to. Sharpen the intuition, not the implementation. He cited Fred Brooks&#8217; &#8220;No Silver Bullet&#8221; and the distinction between essential and incidental complexity. AI is handling a lot of the incidental complexity that accumulated over decades. The essential complexity of building systems that solve real problems for real people has not gone anywhere.</span></p><p><span>&#8220;You&#8217;re just moving further up the stack,&#8221; Webster said, &#8220;and that&#8217;s a skill a lot of folks don&#8217;t have, because for a long time it was really hard to do those other things and now those things are easier.&#8221;</span></p><h2><span>3. College prepared Hanabel for a job that barely exists anymore</span></h2><p><span>Hanabel graduated about a year before joining CircleCI. She came up through a computer science program that treated AI as a form of academic dishonesty: sign at the bottom of every test saying you didn&#8217;t use it, even as everyone quietly used it.</span></p><p><span>The irony Rob raised: using AI would have been cheating. Using AI would have also prepared her for her actual job.</span></p><p><span>Like most fresh graduates, Hanabel found herself equipped with technical skills, but needing experience in everything around them.</span></p><blockquote><p><span>&#8220;I feel like that&#8217;s me during all the meetings. I knew nothing about the process or the delivery side of things. I didn&#8217;t know what that would look like outside of college.&#8221;</span></p><p><strong><span>Hanabel Mengistu, Software Engineer, CircleCI</span></strong></p></blockquote><p><span>Rob&#8217;s take is worth sitting with: &#8220;A lot of the day-to-day job is not actually about the structure of software. It&#8217;s: How do we decide what to work on? How do we communicate about what we&#8217;re working on? How do we decide the risk involved? How do we deal with breakage? That&#8217;s probably more than 50% of the job. And then we took the part we did focus on in school and said, don&#8217;t worry about that anymore.&#8221;</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.confidentcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.confidentcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2><span>4. Juniors aren&#8217;t being mentored, and it&#8217;s showing up as loneliness</span></h2><p><span>One of the quieter observations in this episode: junior engineers are feeling isolated, and it&#8217;s not entirely their fault or anyone else&#8217;s fault specifically. The mentorship pipeline has a structural problem.</span></p><p><span>Hanabel described what she&#8217;s hearing from peers:</span></p><blockquote><p><span>&#8220;A lot of my friends that are in the industry feel a little bit alone. People aren&#8217;t mentoring other juniors anymore because it&#8217;s something AI can do. Not really, but for small questions, you can just ask Claude. And asking AI is not always the most valuable way to learn.&#8221;</span></p><p><strong><span>Hanabel Mengistu, Software Engineer, CircleCI</span></strong></p></blockquote><p><span>The deeper issue Rob named is a two-sided dynamic: senior engineers assume juniors can just ask the model, so they stop offering; junior engineers sense they shouldn&#8217;t bother asking, so they stop reaching out. But the value a good mentor provides isn&#8217;t just the answer to the question. It&#8217;s the context around the question, the adjacent things worth knowing, and the pattern recognition that says &#8220;you&#8217;re asking about X because you&#8217;re missing Y.&#8221;</span></p><h2><span>5. The great leveling might actually help juniors more than it hurts them</span></h2><p><span>The most optimistic thread in this conversation came from an unexpected place. Hanabel pushed back, gently, on the framing that the current moment is particularly hard for new engineers. Her read: it might actually be easier in one specific way.</span></p><blockquote><p><span>&#8220;I feel like it&#8217;s comforting. It&#8217;s nice to be able to figure stuff out with other people and not feel like I&#8217;m the worst. I&#8217;ve been able to contribute so quickly. I wouldn&#8217;t be able to do all the things I&#8217;m contributing to if it weren&#8217;t for that.&#8221;</span></p><p><strong><span>Hanabel Mengistu, Software Engineer, CircleCI</span></strong></p></blockquote><p><span>Claude Code went GA in June 2025. ChatGPT went GA in November 2022. Nobody has ten years of experience in any of this.</span></p><p><span>The experience gap that normally takes years to cross hasn&#8217;t had time to form in the places that matter most for this transition. For junior developers, this means that the playing field has flattened.</span></p><p><span>A curious generalist with no bad habits to break might actually have an edge over a deeply experienced engineer who has spent a decade optimizing for a set of constraints that no longer apply.</span></p><p><em><span>Subscribe to Confident Commit for more conversations about how software delivery is changing, from the people figuring it out in real time: </span><a href="https://confidentcommit.com"><span>confidentcommit.com</span></a></em></p>]]></content:encoded></item><item><title><![CDATA[Guardrails for shipping with AI agents, feat. Luca Rossi of Refactoring.fm]]></title><description><![CDATA[Listen now (31 mins) | Rob Zuber sits down with Luca Rossi, founder of Refactoring.fm, to talk about what high-performing engineering teams are actually getting right with AI in 2026.]]></description><link>https://www.confidentcommit.com/p/guardrails-for-shipping-with-ai-agents-539</link><guid isPermaLink="false">https://www.confidentcommit.com/p/guardrails-for-shipping-with-ai-agents-539</guid><dc:creator><![CDATA[Confident Commit]]></dc:creator><pubDate>Thu, 23 Jul 2026 15:00:00 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/211758774/8dd51775938412590b638e859f401d69.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><a href="https://it.linkedin.com/in/lucaronin"><span>Luca Rossi</span></a><span> has spent five years running </span><a href="https://refactoring.fm/"><span>Refactoring</span></a><span>, a newsletter and community about software engineering read by more than 170,000 engineers and managers. Before that, he spent a decade as a CTO and startup founder. He&#8217;s not a commentator watching AI happen to other people. He&#8217;s building an open source app called Tolaria, writing about it in real time, and using that firsthand experience to ground everything he publishes.</span></p><p><span>Rob sat down with Luca to talk about what high-performing teams are actually getting right with AI in 2026, why code review was always more questionable than we admitted, and how Luca built a three-part framework: guides, gates, and guards, to keep AI agents honest in production code.</span></p><h2><span>1. 2026 is the year of reckoning, not adoption</span></h2><p><span>Last year, AI tooling spread fast. Most teams added something. Some added everything. This year, the question has shifted from &#8220;are we using AI?&#8221; to &#8220;what are we actually getting for it?&#8221;</span></p><p><span>Luca has been watching this across a wide range of teams, and the diversity is striking. There&#8217;s no consensus, no single playbook, and the gap between teams who are extracting real value and teams who are just spending tokens keeps widening.</span></p><blockquote><p><span>&#8220;2025 was more like the year of adoption. Eventually we got to: &#8216;AI is everywhere, 95 plus percent of engineers using AI&#8217;. And this year is more like okay, show me what you&#8217;re getting out of it, as the token spending keeps growing for everybody.&#8221;</span></p><p><strong><span>Luca Rossi, Founder, Refactoring.fm</span></strong></p></blockquote><p><span>The teams pulling ahead are the ones who&#8217;ve gotten specific about where AI actually helps and where the friction still lives.</span></p><h2><span>2. You have to know your SDLC before you can improve it</span></h2><p><span>This sounds obvious. It isn&#8217;t. Luca&#8217;s observation is that the single biggest differentiator between the teams getting results and those who aren&#8217;t is whether engineering leaders actually have visibility into what&#8217;s happening in their delivery process. It&#8217;s about real awareness of where things slow down, where handoffs break, and where decisions stall.</span></p><blockquote><p><span>&#8220;Just being aware of what&#8217;s going on in your developer process. Which means talking with your engineers, knowing where the problems are, where the bottlenecks are, having a good feedback loop where you can actually work on things that you know can be improved.&#8221;</span></p><p><strong><span>Luca Rossi, Founder, Refactoring.fm</span></strong></p></blockquote><p><span>Rob noted that this insight has been true for 30 years. But the AI era has made it more consequential because improvements compound faster when you know where to aim them.</span></p><h2><span>3. Guides, gates, and guards: a framework for trusting AI agents</span></h2><p><span>Luca has been building Tolaria largely without writing code himself, directing AI agents through the full development workflow. That experiment forced him to get precise about trust. The result is a three-part framework he calls guides, gates, and guards.</span></p><p><strong><span>Guides</span></strong><span> are the instructions agents follow: rules in the agents.md file about how code should be written, covering things like TDD, the Boy Scout rule, analytics instrumentation, and localization.</span></p><p><strong><span>Gates</span></strong><span> are deterministic checks that run before anything gets committed. Because agents ignore the guides roughly 10% of the time, you need enforcement that doesn&#8217;t rely on the agent&#8217;s judgment. Luca uses local hooks to check code health, quality, and security and blocks commits that don&#8217;t clear a threshold.</span></p><blockquote><p><span>&#8220;AI can happily ignore these rules every now and then, like ten percent of the time they just don&#8217;t care and do things how they like to do. And so you need deterministic gates to avoid [the fact that] bad things happen in production and get committed and pushed.&#8221;</span></p><p><strong><span>Luca Rossi, Founder, Refactoring.fm</span></strong></p></blockquote><p><strong><span>Guards</span></strong><span> are nightly or weekly audits that catch judgment calls and bigger-picture issues. Did the agent forget to write an architecture decision record? Have performance metrics degraded since the last release? Are there refactoring opportunities the agent couldn&#8217;t see at the task level?</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.confidentcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.confidentcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2><span>4. Velocity without visibility is just faster drift</span></h2><p><span>One of the clearest tensions in the conversation: AI compresses the coding loop, but that doesn&#8217;t automatically speed up software delivery. It can just mean you&#8217;re drowning somewhere else.</span></p><p><span>Luca pointed to code reviews and product planning as the two places teams tend to hit the ceiling. Code gets written faster, but review queues get longer. Requirements get built faster, but upstream work to define what to build doesn&#8217;t scale with the throughput.</span></p><blockquote><p><span>&#8220;We&#8217;re just getting faster at coding, but people are drowning either in code reviews or are too slow at creating new requirements or measuring what&#8217;s been done. So we have to take a holistic view about the whole thing.&#8221;</span></p><p><strong><span>Luca Rossi, Founder, Refactoring.fm</span></strong></p></blockquote><p><span>The teams breaking through aren&#8217;t asking &#8220;how do I make each stage better?&#8221; They&#8217;re asking &#8220;what am I actually trying to get out of this process?&#8221; and redesigning from that question with the tools they now have.</span></p><h2><span>5. Code review was already questionable. Now it&#8217;s untenable.</span></h2><p><span>Luca has been skeptical of traditional code review for longer than AI has been part of the conversation. The asynchronous process was a time sink before agents started generating code at scale. Teams would fight to shave 10 minutes off CI times while PRs sat idle for 18 hours. And the assumption that humans reliably catch important issues by scanning unfamiliar code quickly has always been optimistic.</span></p><p><span>Rob agreed: We were never that excited about reviewing our colleagues&#8217; code. We tolerated it because it was the best available option.</span></p><blockquote><p><span>&#8220;Even if you don&#8217;t account for the quantity, it&#8217;s just a miserable experience. You don&#8217;t want to corner people into spending a lot of their time at work doing miserable things like reviewing AI code.&#8221;</span></p><p><strong><span>Luca Rossi, Founder, Refactoring.fm</span></strong></p></blockquote><p><span>Now that AI agents can generate code far faster than any team can review it, the math is clearly broken. But the answer isn&#8217;t faster humans. It&#8217;s rethinking what &#8220;review&#8221; is actually trying to accomplish and finding better mechanisms for that. Deterministic gates, paired programming for high-stakes changes, and automated quality enforcement upstream before anything reaches a human.</span></p><h2><span>6. Building for yourself is a cheat code with limits</span></h2><p><span>Luca&#8217;s most consistent lens across everything he talked about, including the newsletter, the app, and the AI workflow, is that building something you personally need gives you a feedback loop most teams don&#8217;t have. You know the problem because you&#8217;re living it. You know when the solution works because you feel it.</span></p><p><span>He&#8217;s explicit that his experience building Tolaria is privileged. He&#8217;s working alone, with no legacy code, and can afford to experiment in ways that most real-world teams can&#8217;t. But the underlying principle transfers: the teams getting the most out of AI tools right now are the ones with enough self-knowledge to know what they&#8217;re actually trying to solve.</span></p><p><em><span>Confident Commit is CircleCI&#8217;s podcast for engineering leaders who want to understand how software delivery is actually changing. Subscribe at </span><a href="https://confidentcommit.com"><span>confidentcommit.com</span></a><span> to get every episode plus companion posts like this one.</span></em></p>]]></content:encoded></item><item><title><![CDATA[Untested AI is unshippable AI]]></title><description><![CDATA[Listen now (35 mins) | Rob Zuber chats with Laurie Voss, Head of Developer Relations at Arize to talk about why AI applications ship without real testing, how to build evals that work in CI, and whether agents are ready.]]></description><link>https://www.confidentcommit.com/p/untested-ai-is-unshippable-ai-with-b95</link><guid isPermaLink="false">https://www.confidentcommit.com/p/untested-ai-is-unshippable-ai-with-b95</guid><dc:creator><![CDATA[Confident Commit]]></dc:creator><pubDate>Thu, 18 Jun 2026 13:00:00 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/211758775/245bf4073a2519124516dfea0a984064.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><span>Most AI applications in production right now were shipped on instinct. A developer ran a few favorite queries, liked what they saw, and pushed.</span></p><p><span>Laurie Voss has watched this pattern play out at scale. As co-founder of npm (the package registry that now serves hundreds of millions of JavaScript developers), a DevRel leader at Netlify, and now Head of Developer Relations at Arize, she has spent her career at the intersection of data and developer behavior. She knows how software communities learn, and she knows how they avoid learning.</span></p><p><span>The current moment in AI development, she argues, looks a lot like the early days of software testing: everyone knows tests exist, and most people are not writing them.</span></p><p><span>Rob Zuber, CTO at CircleCI, sat down with Laurie to talk about what &#8220;real evals&#8221; actually are, why the eval loop is going to change how software gets built, and whether we are anywhere close to agents building production software without humans in the loop.</span></p><h2><span>1. Evals are just tests. Stop calling them anything else.</span></h2><p><span>The first problem Laurie has with the word &#8220;evals&#8221; is the word itself. It is borrowed from the world of ML, and it lands with a thud in the ears of engineers who came up through web development, backend systems, or platform work.</span></p><blockquote><p><span>&#8220;As soon as you say traces are logs and evals are tests, AI engineers get the story much better. You&#8217;re supposed to be writing tests.&#8221;</span></p><p><strong><span>Laurie Voss, Head of DevRel, Arize</span></strong></p></blockquote><p><span>The terminology barrier is not trivial. If engineers hear &#8220;evals&#8221; as an ML research concept, it stays over there, in someone else&#8217;s domain. If they hear &#8220;tests,&#8221; the mental model clicks and the excuses disappear. The biggest competition for an evals company, Laurie points out, is not a rival product. It is engineers who just do not do any evals at all.</span></p><p><span>The implication is worth sitting with: a large fraction of AI applications going to production right now have no systematic validation. If you are running evals, you already have an advantage.</span></p><h2><span>2. &#8220;Vibe-based testing&#8221; is why people hate AI.</span></h2><p><span>There is a phrase that stuck: vibe-based testing. It&#8217;s what happens when a developer makes a change, runs a favorite query or two, feels good, and ships. It works fine for the happy path. It does nothing for the other 80% of inputs users will eventually throw at the application.</span></p><blockquote><p><span>&#8220;A whole lot of software is getting all the way to production right now using nothing more than vibes. And it is showing up as really poor reliability. And it is making people hate AI because they interact with AI that&#8217;s busted all the time.&#8221;</span></p><p><strong><span>Laurie Voss, Head of DevRel, Arize</span></strong></p></blockquote><p><span>This is not a new failure mode. Anyone who has watched engineers click-test a UI before a release will recognize the shape of it. The specific damage here is that broken AI applications have the potential to erode trust in AI as a category. The cost is the whole wave of adoption getting slower because people have been burned.</span></p><p><span>If you are building AI applications, the quality of your evals is now a competitive differentiator.</span></p><h2><span>3. LLM-as-judge is the only way to scale, but it needs its own tuning.</span></h2><p><span>Non-determinism is the real wrench. Traditional tests are essentially string matching: you put in an input, you expect an output, you check if it matches. With AI applications, the same input can produce a million different outputs, and a meaningful fraction of them will be correct. You cannot just check for a specific string.</span></p><p><span>For some cases, smarter string matching helps: check for &#8220;one hour ago&#8221; and &#8220;60 minutes ago&#8221; and other valid phrasings. But for anything requiring judgment, you need a judge. In production, at scale, that judge has to be another LLM.</span></p><p><span>The catch is that the judge LLM introduces its own layer of non-determinism. It has its own prompt, and that prompt needs to be tuned against real output to confirm it is actually measuring what you want it to measure, rather than producing good/bad labels at random.</span></p><p><span>If you are running evals in CI, where they need to execute reliably thousands of times a day, this is the architecture you are building toward: a test suite powered by an LLM that has itself been validated as a good judge.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.confidentcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.confidentcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2><span>4. Regression evals vs. capability evals: two different jobs.</span></h2><p><span>One of the most useful frames in the conversation was the distinction between regression evals and capability evals. They are not interchangeable, and conflating them leads to poorly designed test suites.</span></p><p><span>Regression evals are what most people think of first. They are close to traditional tests: can my application still do the things I already know it can do? They are the safety net you run continuously.</span></p><p><span>Capability evals are different in a fundamental way.</span></p><blockquote><p><span>&#8220;Capability evals are a test I expect you to fail. This is a test where I expect you to score 20 percent and then do better the next time and climb a hill of capability.&#8221;</span></p><p><strong><span>Laurie Voss, Head of DevRel, Arize</span></strong></p></blockquote><p><span>The combination opens up serious possibilities. Because evals powered by LLMs include an explanation field alongside a score, you can feed the failure explanations from 100 runs back into a coding agent and ask it to improve the software based on what went wrong. No human required. The eval loop drives the development loop. This has not been fully productionized yet, but the pieces exist.</span></p><h2><span>5. Context engineering is the whole game.</span></h2><p><span>The conversation shifted toward agentic coding and the question of how close we are to agents that can build production software end to end. Laurie&#8217;s answer was grounded in a single concept: context engineering.</span></p><p><span>Every capability question in AI development comes back to context. Can this LLM be a doctor? That requires a human lifetime of context, and it does not fit in a window. Can this LLM build a website? Websites are a narrow enough domain that Laurie thinks we are close to an agent harness that can handle it reliably. Can this LLM build production software in general? Software development as a domain is too large for any current window or graph structure to cover well enough to one-shot it.</span></p><p><span>Context graphs, one of the tools Arize is working on, are a form of context compression: a way of defining which topics connect to which other topics so an LLM can navigate a large domain without losing things out the bottom of its window.</span></p><p><span>The practical takeaway for engineering teams building on AI today: the investment in skills files, harnesses, and structured context is what determines whether your agent makes good decisions or system-level mistakes.</span></p><h2><span>6. Software developers always solve their own problems first.</span></h2><p><span>There is a pattern in software development that Laurie named plainly: the community perfects tooling for its own domain before solving anyone else&#8217;s. The reason there are so many frameworks is that developers made their own environments excellent before turning to enterprise software, which is still, by most measures, remarkably bad.</span></p><p><span>The same dynamic is happening now. AI coding tools are improving fast because developers are building them and using them on their own work. The patterns being developed in AI coding evals, harness design, and context engineering will eventually transfer to other domains. But that transfer comes later.</span></p><p><span>For teams building AI products now, the implication is to pay attention to what the software development domain is learning. The patterns being forged here will be the foundations for whatever comes next.</span></p><p><em><span>If this episode got you nodding your head, subscribe to </span><a href="https://www.confidentcommit.com/"><span>Confident Commit</span></a><span> for more conversations like this one.</span></em></p>]]></content:encoded></item><item><title><![CDATA[What snake games have taught us about shipping with AI agents]]></title><description><![CDATA[Five months of Loop Lab Snake builds trace a path from AFK Ralph loops to 100% green PRs and 3x faster CI feedback. Same benchmark, four breakthroughs.]]></description><link>https://www.confidentcommit.com/p/what-snake-games-have-taught-us-about</link><guid isPermaLink="false">https://www.confidentcommit.com/p/what-snake-games-have-taught-us-about</guid><dc:creator><![CDATA[Ryan E. Hamilton]]></dc:creator><pubDate>Wed, 10 Jun 2026 21:39:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/43bc3749-17a2-4672-856e-5ff7eff0b2fc_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Back in January, I wired up a </span><a href="https://ghuntley.com/loop/"><span>Ralph loop</span></a><span> to build Snake games and started dabbling in AFK development, with appropriate levels of human supervision, as needed.</span></p><p><span>Five months later, I am still coding the same game of Snake. Now 100% AFK, with much more confidence than when I started. Zero human supervision.</span></p><p><span>The first half of 2026 has been fun, but not in the arcade sense. It has been fun in the </span><strong><span>&#8220;I haven&#8217;t touched my keyboard in twenty minutes, and every check on the branch is green&#8221;</span></strong><span> sense. We kept building the same Snake game. Same 20x20 grid. Same arrow keys. Same TDD spec. Same retro aesthetic.</span></p><p><strong><span>The game has barely changed, but the delivery loop has changed everything.</span></strong></p><p><span>If you have been watching the AI agent hype cycle from a safe distance, here is the compressed version of what we learned by staying stubbornly on one benchmark: </span><strong><span>you can go AFK with rising confidence, greener pull requests, and faster feedback than we have ever measured. Not by trusting the agent more. By giving it better and quicker back-pressure at each layer of validation.</span></strong></p><h2><span>The benchmark nobody asked for (but everybody needed)</span></h2><p><span>Snake is a toy. That is the point.</span></p><p><span>A toy spec is small enough to run ten times in a week. Small enough to isolate one variable at a time. Small enough that when something breaks, you can actually read the diff instead of drowning in a monorepo.</span></p><p><span>We gave agents the same task repeatedly: build a playable Snake game from scratch using test-driven development (TDD). Seven tasks. Canvas rendering. Collision detection. Score tracking. Push commits. Open a PR. Walk away.</span></p><p><span>Every run finished locally. Every agent said the tests passed. Every game worked on the laptop and was pretty fun to play.</span></p><h2><span>January: everything is a Ralph loop</span></h2><p><span>In January, Geoffrey Huntley published </span><a href="https://ghuntley.com/loop/"><span>everything is a ralph loop</span></a><span>. The mindset shift is not &#8220;use AI to type faster.&#8221; It is &#8220;program the loop.&#8221;</span></p><p><span>Allocate a goal. Run the loop. Watch the failures. Fix the failure domain so it never happens again. Repeat until done or until you hit CTRL+C and take the wheel back.</span></p><p><span>Basic Ralph loops can build software and generate PRs AFK. That part worked decently well (with some light human supervision).</span></p><p><strong><span>The PR at the end might be green. It might be red. The agent does not inherently know which until something external tells it.</span></strong></p><p><span>We were building Snake games in loops before we had a name for what we were doing.</span></p><p><span>The lesson from January: </span><strong><span>autonomy without validation is just faster guessing.</span></strong></p><h2><span>February: does the pipeline agree with local tests?</span></h2><p><span>By February we had data, not vibes, on whether the pipeline agreed.</span></p><p><a href="https://loop.circleci.com/we-let-an-ai-agent-say-i-passed-was-it-actually-good"><span>Our first study</span></a><span> ran ten controlled Snake builds. Same spec. Same model scaffolding. One variable: five runs wired to CircleCI through </span><a href="https://github.com/CircleCI-Research/ralph-ci"><span>RalphCI</span></a><span>, five runs without.</span></p><p><span>All ten agents completed every task. All ten passed local tests. All ten declared victory.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Ntek!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa53aec79-ecb9-4d0d-adec-ad1fc6283b30_1292x344.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Ntek!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa53aec79-ecb9-4d0d-adec-ad1fc6283b30_1292x344.png 424w, https://substackcdn.com/image/fetch/$s_!Ntek!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa53aec79-ecb9-4d0d-adec-ad1fc6283b30_1292x344.png 848w, https://substackcdn.com/image/fetch/$s_!Ntek!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa53aec79-ecb9-4d0d-adec-ad1fc6283b30_1292x344.png 1272w, https://substackcdn.com/image/fetch/$s_!Ntek!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa53aec79-ecb9-4d0d-adec-ad1fc6283b30_1292x344.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Ntek!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa53aec79-ecb9-4d0d-adec-ad1fc6283b30_1292x344.png" width="1292" height="344" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/a53aec79-ecb9-4d0d-adec-ad1fc6283b30_1292x344.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:344,&quot;width&quot;:1292,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:40760,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/212608969?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa53aec79-ecb9-4d0d-adec-ad1fc6283b30_1292x344.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Ntek!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa53aec79-ecb9-4d0d-adec-ad1fc6283b30_1292x344.png 424w, https://substackcdn.com/image/fetch/$s_!Ntek!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa53aec79-ecb9-4d0d-adec-ad1fc6283b30_1292x344.png 848w, https://substackcdn.com/image/fetch/$s_!Ntek!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa53aec79-ecb9-4d0d-adec-ad1fc6283b30_1292x344.png 1272w, https://substackcdn.com/image/fetch/$s_!Ntek!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fa53aec79-ecb9-4d0d-adec-ad1fc6283b30_1292x344.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Eighty percent of the no-CI runs shipped code the pipeline rejected. The failure was mundane: single quotes where ESLint wanted double quotes. Local tests do not check lint rules. CI does. The gap between &#8220;works on my machine&#8221; and &#8220;works in your pipeline&#8221; is where agents lie with confidence.</span></p><p><span>RalphCI fixed that by injecting real pipeline status into the agent loop after push. When CI failed, a CI Doctor agent read the logs, applied fixes, and pushed again. Twelve failures across five runs. Twelve autonomous fixes. Zero human triage.</span></p><p><span>The breakthrough in February: </span><strong><span>agentic loops with CI back-pressure can land green PRs AFK.</span></strong><span> The final PR passes CI. Commits inside that PR might still fail along the way. Red commits on the branch. Green at the end. Progress, not perfection.</span></p><h2><span>Early May: every commit has to be green</span></h2><p><span>Three months later we hit the next failure domain.</span></p><p><span>Local pnpm test:run kept passing on my Mac. </span><a href="https://circleci.com/blog/chunk-sidecars/"><span>Chunk sidecars</span></a><span> kept failing on Linux. Missing LEGAL_DISCLAIMER.md. Missing SNAKE.md experiment marker. A TypeScript compile error the Snake workspace&#8217;s local tests never exercised but a full pnpm build caught.</span></p><p><span>Same pattern, five runs in a row. Local green. Sidecar red. CI Doctor lands a </span><code>fix(ci-sidecar):</code><span> commit. Sidecar green. Push allowed. </span></p><p><a href="https://loop.circleci.com/afk-builds-with-100-green-prs-chunk-sidecars-inside-the-agent-loop"><span>Our early May study</span></a><span> wired Chunk sidecars into RalphCI&#8217;s Review Gate. A sidecar is a lightweight remote microVM: your workspace syncs there, Chunk runs a microbuild that mirrors your pipeline intent, and you get CI-shaped feedback before the branch leaves your laptop.</span></p><p><strong><span>Five Snake builds. Five 100% green PRs. Every commit green. Not just the last one.</span></strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vnzw!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffed4a882-9898-4a8a-aef2-35a59e7f972a_1304x474.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vnzw!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffed4a882-9898-4a8a-aef2-35a59e7f972a_1304x474.png 424w, https://substackcdn.com/image/fetch/$s_!vnzw!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffed4a882-9898-4a8a-aef2-35a59e7f972a_1304x474.png 848w, https://substackcdn.com/image/fetch/$s_!vnzw!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffed4a882-9898-4a8a-aef2-35a59e7f972a_1304x474.png 1272w, https://substackcdn.com/image/fetch/$s_!vnzw!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffed4a882-9898-4a8a-aef2-35a59e7f972a_1304x474.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vnzw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffed4a882-9898-4a8a-aef2-35a59e7f972a_1304x474.png" width="1304" height="474" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fed4a882-9898-4a8a-aef2-35a59e7f972a_1304x474.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:474,&quot;width&quot;:1304,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:62245,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/212608969?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffed4a882-9898-4a8a-aef2-35a59e7f972a_1304x474.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vnzw!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffed4a882-9898-4a8a-aef2-35a59e7f972a_1304x474.png 424w, https://substackcdn.com/image/fetch/$s_!vnzw!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffed4a882-9898-4a8a-aef2-35a59e7f972a_1304x474.png 848w, https://substackcdn.com/image/fetch/$s_!vnzw!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffed4a882-9898-4a8a-aef2-35a59e7f972a_1304x474.png 1272w, https://substackcdn.com/image/fetch/$s_!vnzw!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffed4a882-9898-4a8a-aef2-35a59e7f972a_1304x474.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>I had not touched the keyboard in twenty minutes. The final PR (and every commit on it) was green anyway.</span></p><p><strong><span>This screenshot is what HIGH CONFIDENCE looks like.</span></strong></p><p><span>The May breakthrough: </span><strong><span>three validation layers (local, sidecar, pipeline) produce PRs where humans review green branches only.</span></strong><span> Sidecars catch CI-shaped failures while the agent still has context. CircleCI remains the authority at merge time. </span><strong><span>Sidecars extend CI into the inner loop. They do not replace it.</span></strong></p><h2><span>Late May: deliver the same green outcome, faster</span></h2><p><span>100% Green PRs answered &#8220;can we merge with confidence?&#8221; A follow-up experiment asked &#8220;how fast can we merge with confidence?&#8221;</span></p><p><a href="https://loop.circleci.com/the-sidecar-race-22-seconds-vs-69-seconds-inside-the-agent-loop"><span>Our late May study</span></a><span> ran the same ten Go tasks two ways on </span><code>chunk-cli:</code><span> sidecar validation per iteration versus commit-push-poll CircleCI per iteration. Same agent. Same model. Same gate jobs </span><code>lint</code><span> and </span><code>test</code><span>). Different machinery.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XYxh!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a0183c3-7fe3-4dd3-b150-c6ab5d06504a_1292x328.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XYxh!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a0183c3-7fe3-4dd3-b150-c6ab5d06504a_1292x328.png 424w, https://substackcdn.com/image/fetch/$s_!XYxh!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a0183c3-7fe3-4dd3-b150-c6ab5d06504a_1292x328.png 848w, https://substackcdn.com/image/fetch/$s_!XYxh!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a0183c3-7fe3-4dd3-b150-c6ab5d06504a_1292x328.png 1272w, https://substackcdn.com/image/fetch/$s_!XYxh!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a0183c3-7fe3-4dd3-b150-c6ab5d06504a_1292x328.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XYxh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a0183c3-7fe3-4dd3-b150-c6ab5d06504a_1292x328.png" width="1292" height="328" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/0a0183c3-7fe3-4dd3-b150-c6ab5d06504a_1292x328.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:328,&quot;width&quot;:1292,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:46259,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/212608969?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a0183c3-7fe3-4dd3-b150-c6ab5d06504a_1292x328.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!XYxh!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a0183c3-7fe3-4dd3-b150-c6ab5d06504a_1292x328.png 424w, https://substackcdn.com/image/fetch/$s_!XYxh!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a0183c3-7fe3-4dd3-b150-c6ab5d06504a_1292x328.png 848w, https://substackcdn.com/image/fetch/$s_!XYxh!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a0183c3-7fe3-4dd3-b150-c6ab5d06504a_1292x328.png 1272w, https://substackcdn.com/image/fetch/$s_!XYxh!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F0a0183c3-7fe3-4dd3-b150-c6ab5d06504a_1292x328.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Sixty-nine seconds. Every iteration. Five replicates in a row. That is not random noise in our harness. That is the queue tax on push-per-task CI.</span></p><p><span>Twenty-two seconds on a warmed sidecar snapshot. Same question: are lint and test happy right now? Same answer. Different delivery speed.</span></p><p><span>Token spend was flat because in this particular experiment tokens tracked fixes, not wait. The CI arm burned more clock time, not more intelligence, so sidecars did not shrink the LLM bill here. They shrunk idle time. On a ten-task run that is roughly eight minutes of gate-waiting saved before the final pipeline epilogue.</span></p><p><span>The May speed breakthrough: </span><strong><span>you still want an outer loop.</span></strong><span> Sidecar runs finished with a full ci workflow epilogue after one push. Fast inner loop on each task, authoritative pipeline confirmation at the end. Not sidecar-only cowboy coding. Not push-and-pray either.</span></p><h2><span>Four layers, one arc</span></h2><p><span>Strip away the Snake skin and the arc looks like this:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!szuq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75991a28-536e-4f63-92f1-79661f53e883_1284x636.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!szuq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75991a28-536e-4f63-92f1-79661f53e883_1284x636.png 424w, https://substackcdn.com/image/fetch/$s_!szuq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75991a28-536e-4f63-92f1-79661f53e883_1284x636.png 848w, https://substackcdn.com/image/fetch/$s_!szuq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75991a28-536e-4f63-92f1-79661f53e883_1284x636.png 1272w, https://substackcdn.com/image/fetch/$s_!szuq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75991a28-536e-4f63-92f1-79661f53e883_1284x636.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!szuq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75991a28-536e-4f63-92f1-79661f53e883_1284x636.png" width="1284" height="636" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/75991a28-536e-4f63-92f1-79661f53e883_1284x636.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:636,&quot;width&quot;:1284,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:123775,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/212608969?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75991a28-536e-4f63-92f1-79661f53e883_1284x636.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!szuq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75991a28-536e-4f63-92f1-79661f53e883_1284x636.png 424w, https://substackcdn.com/image/fetch/$s_!szuq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75991a28-536e-4f63-92f1-79661f53e883_1284x636.png 848w, https://substackcdn.com/image/fetch/$s_!szuq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75991a28-536e-4f63-92f1-79661f53e883_1284x636.png 1272w, https://substackcdn.com/image/fetch/$s_!szuq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F75991a28-536e-4f63-92f1-79661f53e883_1284x636.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>We are not gaining more AFK confidence because we trust agents more (although the models have improved since January). We are gaining more AFK confidence because each failure domain we hit got engineered out of the loop. And we&#8217;ve even unleashed some speed gains along the way.</span></p><p><strong><span>Watch the loop.</span></strong><span> That is </span><a href="https://ghuntley.com/loop/"><span>Huntley&#8217;s line</span></a><span> and it is the whole game.</span></p><ul><li><p><a href="https://ghuntley.com/loop/"><span>January</span></a><span> taught us to program the AFK agent loop.</span></p></li><li><p><a href="https://loop.circleci.com/we-let-an-ai-agent-say-i-passed-was-it-actually-good"><span>February</span></a><span> taught us the agent&#8217;s &#8220;I passed&#8221; is local theater without CI.</span></p></li><li><p><a href="https://loop.circleci.com/afk-builds-with-100-green-prs-chunk-sidecars-inside-the-agent-loop"><span>Early May</span></a><span> taught us that 100% green PRs are possible, given sidecar validation before commit and push.</span></p></li><li><p><a href="https://loop.circleci.com/the-sidecar-race-22-seconds-vs-69-seconds-inside-the-agent-loop"><span>Late May</span></a><span> taught us that 100% green PRs are still the prize, and sidecars get us there faster than push-per-task CI.</span></p></li></ul><p><span>Snake games did not teach us that AI writes good code. They taught us </span><strong><span>where validation has to live</span></strong><span> when the coder is an agent loop and the reviewer is still a human with a merge button.</span></p><h2><span>What I am actually doing with this</span></h2><p><span>I am obviously not shipping Snake to production. (Unless you&#8217;ve been craving retro gaming experiences. </span><a href="https://www.linkedin.com/in/ryanehamilton/"><span>DM me.</span></a><span>)</span></p><p><span>I am shipping the stack: </span><a href="https://github.com/CircleCI-Research/ralph-ci"><span>RalphCI</span></a><span> orchestration, Review Gate with local checks plus Chunk sidecar validation, CI Doctor for autonomous repair, CircleCI as the outer loop authority. Same pattern we pressure-tested on a toy game because toys fail fast and teach faster.</span></p><p><span>Four layers, one arc. Program the loop. Wire in CI back-pressure. Validate on sidecars before push. Maintain CI pipeline authority at the end.</span></p><p><span>Each layer kills a failure domain the last one missed. Confidence climbs. PRs go greener. Feedback gets faster. The human at the merge button only reviews green.</span></p><p><span>That is what five months of the same Snake game bought us.</span></p><p><strong><span>Green CI is still priceless. Everything else is noise.</span></strong></p><h2><span>Where to read the lab reports</span></h2><ul><li><p><a href="https://ghuntley.com/loop/"><span>everything is a ralph loop</span></a><span> (Geoffrey Huntley, Jan 2026)</span></p></li><li><p><a href="https://loop.circleci.com/we-let-an-ai-agent-say-i-passed-was-it-actually-good"><span>We Let an AI Agent Say &#8220;I Passed.&#8221; Was It Actually Good?</span></a><span> (Feb 2026)</span></p></li><li><p><a href="https://loop.circleci.com/afk-builds-with-100-green-prs-chunk-sidecars-inside-the-agent-loop"><span>AFK Builds with 100% Green PRs</span></a><span> (May 2026)</span></p></li><li><p><a href="https://loop.circleci.com/the-sidecar-race-22-seconds-vs-69-seconds-inside-the-agent-loop"><span>The Sidecar Race: 22 Seconds vs 69 Seconds</span></a><span> (May 2026)</span></p></li></ul><p><span>More Snake runs coming. Heavier repos. Flakier tests. Parallel agents. Improved CLI. Better agent-facing APIs. You name it.</span></p><p><strong><span>The benchmark stays small on purpose, and we intend to scale the lessons as we go.</span></strong></p><p></p><h4></h4>]]></content:encoded></item><item><title><![CDATA[You can't review code you've never written]]></title><description><![CDATA[Listen now (44 mins) | Rob Zuber sits down with Hywel Carver, founder and CEO of Skiller Whale, to talk about what it actually takes to grow engineers when AI is generating most of the code.]]></description><link>https://www.confidentcommit.com/p/the-skill-gap-ai-cant-close-with-c16</link><guid isPermaLink="false">https://www.confidentcommit.com/p/the-skill-gap-ai-cant-close-with-c16</guid><dc:creator><![CDATA[Confident Commit]]></dc:creator><pubDate>Thu, 04 Jun 2026 13:00:00 GMT</pubDate><enclosure url="https://api.substack.com/feed/podcast/211758776/30e2f70766075f084b68e6e9e3662724.mp3" length="0" type="audio/mpeg"/><content:encoded><![CDATA[<p><span>There&#8217;s a version of the AI coding conversation that goes like this: AI writes the code, humans review it, everyone ships faster. What that version skips is the uncomfortable question underneath: can you review code you&#8217;ve never written yourself?</span></p><p><a href="https://uk.linkedin.com/in/hywelc"><span>Hywel Carver</span></a><span> has been thinking about this longer than most. He&#8217;s the founder and CEO of </span><a href="https://skillerwhale.com/"><span>Skiller Whale</span></a><span>, a company built on the premise that engineers learn best by doing, not by watching. When he first talked publicly about what AI-generated code would mean for developer skill development, it was 2023, at a conference. The consensus at the time was still pretty comfortable. Things have gotten less comfortable since.</span></p><p><span>Rob sat down with Hywel to work through what it actually means to grow engineers in a world where most code is generated, not written.</span></p><h2><span>1. Knowledge transfer is not the same thing as learning</span></h2><p><span>The way most organizations approach skill development is to throw information at people and call it training. Watch this video. Read this documentation. Here&#8217;s a course with 47 modules.</span></p><blockquote><p><span>&#8220;Very often the way that we expect people to learn is by just giving them a load of information. The actual practice of trying to do that thing, we leave you on your own to do in isolation. And that&#8217;s where the rubber hits the road. That&#8217;s when you find out whether you actually can do it or not.&#8221;</span></p><p><strong><span>Hywel Carver, Founder and CEO, Skiller Whale</span></strong></p></blockquote><p><span>The problem is that information without practice doesn&#8217;t produce skill. Hywel&#8217;s framing: knowing and understanding are the first two rungs of a six-level learning hierarchy. Actually being able to apply something, evaluate someone else&#8217;s work, and eventually create something new comes later, and it only gets there through doing. If your development program stops at the second rung and calls it done, you haven&#8217;t built anything.</span></p><h2><span>2. Bloom&#8217;s taxonomy explains exactly why AI code review is harder than it looks</span></h2><p><span>Hywel referenced Bloom&#8217;s taxonomy of cognitive learning as a framework for understanding the current problem. The six levels build on each other: know, understand, apply, analyze, evaluate, create.</span></p><p><span>The way engineers have historically gotten good at reviewing code is by writing a lot of it themselves and having it reviewed in return. That cycle is how you build the mental model that lets you look at a loop and immediately think about off-by-one errors and variable state, rather than reading each line individually.</span></p><blockquote><p><span>&#8220;We are used to a world where applying the skill of writing software is how we get good enough to evaluate it, to look at what someone else has done and say, yes, I am happy to put my name to that and for it to go into production.&#8221;</span></p><p><strong><span>Hywel Carver, Founder and CEO, Skiller Whale</span></strong></p></blockquote><p><span>If 90% of code is now AI-generated, the training loop that used to happen organically through writing code no longer fires at the same rate. You can still develop those mental models. You just have to do it deliberately, through structured practice designed to build them, rather than picking them up as a byproduct of shipping.</span></p><h2><span>3. AI and humans are bad at different things, and that&#8217;s precisely why both matter</span></h2><p><span>Another interesting argument from the conversation: the error categories that AI is susceptible to are largely not the ones humans fall into, and the reverse is also true. AI is excellent at catching syntax errors, off-by-one issues, and assignment-vs-comparison bugs. Humans are not particularly good at those.</span></p><p><span>Humans are good at something different: reasoning about real-world consequences that weren&#8217;t in the scope of the ticket. The change that would give a specific category of users infinite free access. The knock-on effect that shows up in a workflow nobody thought to include in the spec.</span></p><blockquote><p><span>&#8220;We are good at detecting things like: this is going to have a knock-on effect to that category of user that wasn&#8217;t in the scope of the change request. Humans are good at spotting that kind of thing.&#8221;</span></p><p><strong><span>Hywel Carver, Founder and CEO, Skiller Whale</span></strong></p></blockquote><p><span>Hywel presented an example of an AI agent that found production credentials in a repo and wiped the database in nine seconds. It&#8217;s an argument for keeping human and AI review in parallel, because they&#8217;re checking for fundamentally different things.</span></p><p class="button-wrapper" data-attrs="{&quot;url&quot;:&quot;https://www.confidentcommit.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}" data-component-name="ButtonCreateButton"><a class="button primary" href="https://www.confidentcommit.com/subscribe?"><span>Subscribe now</span></a></p><h2><span>4. Using AI well is a skill, and it&#8217;s learnable</span></h2><p><span>There&#8217;s an assumption baked into a lot of AI tooling rollouts: give people access to the tool, maybe run a vendor demo, and let adoption happen. Skiller Whale ran a study that suggests this assumption is wrong.</span></p><p><span>They worked with part of an engineering organization while the rest of the org had access to the same tools but no structured training. Both groups had the vendor talk. Both had access to the AI coding assistant.</span></p><blockquote><p><span>&#8220;The group we worked with had 90% weekly active users where the rest of the org had 50% weekly active users. When we looked at their main measure of productivity, which was PR throughput, the group we worked with saw a roughly doubling in their PR throughput. The rest of the org actually went down very slightly.&#8221;</span></p><p><strong><span>Hywel Carver, Founder and CEO, Skiller Whale</span></strong></p></blockquote><p><span>A doubling versus a slight decline, same tools, same company. The difference was whether engineers had built mental models for how the tools actually work. People who understand what an LLM is doing underneath prompt differently, catch failures differently, and know what to try when the output is wrong.</span></p><h2><span>5. Mental models don&#8217;t expire, even when implementations do</span></h2><p><span>The pace of change in AI tooling feels genuinely overwhelming if you try to keep up at the surface level: new models, new frameworks, new prompting techniques, new papers every week. Hywel&#8217;s argument is that this is the wrong level to track.</span></p><blockquote><p><span>&#8220;If you have a good sense of how an LLM works, and then someone says, one of the things an LLM can do is produce output that says do a function call, you see how tools work and you&#8217;re like, okay, that makes sense to me. You have a way better understanding of it than someone who&#8217;s just been a user of those systems.&#8221;</span></p><p><strong><span>Hywel Carver, Founder and CEO, Skiller Whale</span></strong></p></blockquote><p><span>His analogy: if you had a solid mental model of HTML, CSS, JavaScript, and how a browser renders pages in the mid-1990s, that model has not meaningfully changed today. The implementations are faster, the APIs are cleaner, but the underlying structure is the same. The mental model carries you through decades of surface-level change. The same logic applies to LLMs. The tips-and-tricks content will be obsolete in six months. But the architecture of how they&#8217;re trained and how they produce output hasn&#8217;t changed that much.</span></p><h2><span>6. The job title is changing faster than anyone expected</span></h2><p><span>One thing Hywel raised that didn&#8217;t get a full chapter in the conversation but probably deserves one: after teams go through structured AI training, the next thing they come back asking for is communication skills. Engineers are having more contact with stakeholders. Product functions at some companies have gotten smaller. Individual engineers are doing more.</span></p><p><span>That sounds like good news for engineering influence. It also means engineers who were hired to write code are now expected to run conversations they weren&#8217;t trained for, translate technical decisions for non-technical audiences, and navigate organizational dynamics that used to be someone else&#8217;s job.</span></p><p><span>The skill set that made someone an excellent engineer in 2020 is still valuable, but it needs some expansion in order to meet the role&#8217;s new level of complexity.</span></p><p><em><span>Subscribe to Confident Commit for new episodes and the data behind how software actually ships: </span><a href="https://confidentcommit.com"><span>confidentcommit.com</span></a><span>.</span></em></p>]]></content:encoded></item><item><title><![CDATA[The sidecar race: 22 seconds vs 69 seconds inside the agent loop]]></title><description><![CDATA[A controlled A/B in chunk-cli: same lint+test gates, sidecar remote validate vs push-per-task CI. Median time to signal 3.1x faster on sidecar; LLM costs relatively flat.]]></description><link>https://www.confidentcommit.com/p/the-sidecar-race-22-seconds-vs-69</link><guid isPermaLink="false">https://www.confidentcommit.com/p/the-sidecar-race-22-seconds-vs-69</guid><dc:creator><![CDATA[Ryan E. Hamilton]]></dc:creator><pubDate>Thu, 28 May 2026 21:38:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/f0e18f88-a5df-4c61-8ef3-ba21a7e423d2_1300x830.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong><span>Sixty-nine seconds.</span></strong> That is how long the agent waited, again, for CircleCI to answer a question it had already asked nine times that run: did <code>lint</code> and <code>test</code> pass?</p><p><strong><span>Twenty-two seconds</span></strong><span> is what the other arm averaged for the same question on a </span><a href="https://circleci.com/blog/chunk-sidecars/"><span>Chunk sidecar</span></a><span>.</span></p><p><span>Same repo. Same ten Go tasks. Same Claude Agent SDK edits. Same gate jobs. Different loop.</span></p><p><span>Three weeks ago I published </span><a href="https://www.confidentcommit.com/p/afk-builds-with-100-green-prs-chunk"><span>AFK Builds with 100% Green PRs</span></a><span>. Five Snake games. </span><a href="https://loop.circleci.com/hardening-ralphci-loops-for-open-source-after-the-february-2026-study"><span>RalphCI</span></a><span> orchestration. Sidecars in the Review Gate so agents stopped feeding red commits to GitHub. </span><strong><span>Green PRs were the win.</span></strong></p><p><span>This follow-up is narrower and, I think, more revealing for anyone running agents today: </span><strong><span>not whether sidecars help you merge green, but how fast the agent gets CI-shaped feedback while the file is still open.</span></strong></p><p><span>I kicked the tires on sidecars again.</span></p><p><strong><span>They are fast.</span></strong></p><p><span>Not a much cheaper agent. A less idle one. Token spend was flat in this particular experiment (more on that later).</span></p><p><strong><span>The win was clock time.</span></strong></p><h2><span>Hypothesis</span></h2><p><span>If each agent task is validated with </span><code>chunk sidecar sync</code><span data-color="#6aa84f" style="color: rgb(106, 168, 79);"> </span><span>plus </span><code>chunk validate --remote</code><span> </span><code>lint</code><span> and </span><code>test-changed</code><span>) instead of </span><strong><span>commit &#8594; push &#8594; poll CircleCI</span></strong><span> for the same </span><code>lint</code><span data-color="#6aa84f" style="color: rgb(106, 168, 79);"> </span><span>and </span><code>test</code><span> jobs, then:</span></p><ol><li><p><strong><span>Median time to signal (TTS)</span></strong><span> per iteration will be materially lower on the sidecar arm.</span></p></li><li><p><strong><span>LLM cost</span></strong><span> will be similar (same agent, same ten prompts, same model fixing the same mistakes).</span></p></li><li><p><strong><span>CircleCI still matters</span></strong><span> for pipeline-level confirmation (sidecar runs end with a full </span><code>ci</code><span> workflow epilogue after one final push).</span></p></li></ol><p><span>I expected push-per-task CI to pay a queue tax every iteration. However, I did not expect the median to land on </span><strong><span>69 seconds five replicates in a row</span></strong><span>.</span></p><h2><span>Setup</span></h2><p><strong><span>Repo:</span></strong></p><ul><li><p><a href="https://github.com/CircleCI-Public/chunk-cli"><span>CircleCI-Public/chunk-cli</span></a></p></li><li><p><a href="https://github.com/CircleCI-Public/chunk-cli/pull/370"><span>Experiment PR #370</span></a></p></li></ul><p><strong><span>Tasks:</span></strong><span> Ten cumulative edits on a small Go fixture </span><code>internal/racefixture/</code><span> on experimental run branches).</span></p><p><strong><span>Replicates:</span></strong><span> Five labels </span><code>001&#8211;005</code><span>) per arm. </span><strong><span>10 recorded runs total.</span></strong></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Xocf!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7922abaf-06d7-486a-a570-52da64e589f8_1294x540.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Xocf!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7922abaf-06d7-486a-a570-52da64e589f8_1294x540.png 424w, https://substackcdn.com/image/fetch/$s_!Xocf!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7922abaf-06d7-486a-a570-52da64e589f8_1294x540.png 848w, https://substackcdn.com/image/fetch/$s_!Xocf!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7922abaf-06d7-486a-a570-52da64e589f8_1294x540.png 1272w, https://substackcdn.com/image/fetch/$s_!Xocf!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7922abaf-06d7-486a-a570-52da64e589f8_1294x540.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Xocf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7922abaf-06d7-486a-a570-52da64e589f8_1294x540.png" width="1294" height="540" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/7922abaf-06d7-486a-a570-52da64e589f8_1294x540.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:540,&quot;width&quot;:1294,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:89884,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/212614385?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7922abaf-06d7-486a-a570-52da64e589f8_1294x540.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Xocf!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7922abaf-06d7-486a-a570-52da64e589f8_1294x540.png 424w, https://substackcdn.com/image/fetch/$s_!Xocf!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7922abaf-06d7-486a-a570-52da64e589f8_1294x540.png 848w, https://substackcdn.com/image/fetch/$s_!Xocf!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7922abaf-06d7-486a-a570-52da64e589f8_1294x540.png 1272w, https://substackcdn.com/image/fetch/$s_!Xocf!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F7922abaf-06d7-486a-a570-52da64e589f8_1294x540.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><strong><span>Metric:</span></strong><span> </span><strong><span>TTS</span></strong><span> = wall-clock seconds from iteration start until both gates report pass/fail (logged in </span><code>results.csv</code>`<span>).</span></p><p><span>CircleCI CTO </span><a href="https://circleci.com/blog/chunk-sidecars/"><span>Rob Zuber</span></a><span> frames this as </span><strong><span>rebalancing inner and outer loop validation</span></strong><span>. My </span><a href="https://loop.circleci.com/afk-builds-with-100-green-prs-chunk-sidecars-inside-the-agent-loop"><span>May post</span></a><span> showed the outcome layer (green PRs). This post measures the </span><strong><span>wait layer</span></strong><span>.</span></p><h2><span>Results</span></h2><h3><span>Headline: median time to signal</span></h3><p><span>Aggregate: </span><strong><span>median of per-run medians</span></strong><span> across five replicates.</span></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!44NO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F827048fb-b286-404e-8cdf-4b0b8b823726_1286x306.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!44NO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F827048fb-b286-404e-8cdf-4b0b8b823726_1286x306.png 424w, https://substackcdn.com/image/fetch/$s_!44NO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F827048fb-b286-404e-8cdf-4b0b8b823726_1286x306.png 848w, https://substackcdn.com/image/fetch/$s_!44NO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F827048fb-b286-404e-8cdf-4b0b8b823726_1286x306.png 1272w, https://substackcdn.com/image/fetch/$s_!44NO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F827048fb-b286-404e-8cdf-4b0b8b823726_1286x306.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!44NO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F827048fb-b286-404e-8cdf-4b0b8b823726_1286x306.png" width="1286" height="306" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/827048fb-b286-404e-8cdf-4b0b8b823726_1286x306.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:306,&quot;width&quot;:1286,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:40253,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/212614385?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F827048fb-b286-404e-8cdf-4b0b8b823726_1286x306.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!44NO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F827048fb-b286-404e-8cdf-4b0b8b823726_1286x306.png 424w, https://substackcdn.com/image/fetch/$s_!44NO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F827048fb-b286-404e-8cdf-4b0b8b823726_1286x306.png 848w, https://substackcdn.com/image/fetch/$s_!44NO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F827048fb-b286-404e-8cdf-4b0b8b823726_1286x306.png 1272w, https://substackcdn.com/image/fetch/$s_!44NO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F827048fb-b286-404e-8cdf-4b0b8b823726_1286x306.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><h3>Five replicates at a glance</h3><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!HogY!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c0a0a62-6ac0-4469-946d-d3c8e636b7c9_1294x468.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!HogY!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c0a0a62-6ac0-4469-946d-d3c8e636b7c9_1294x468.png 424w, https://substackcdn.com/image/fetch/$s_!HogY!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c0a0a62-6ac0-4469-946d-d3c8e636b7c9_1294x468.png 848w, https://substackcdn.com/image/fetch/$s_!HogY!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c0a0a62-6ac0-4469-946d-d3c8e636b7c9_1294x468.png 1272w, https://substackcdn.com/image/fetch/$s_!HogY!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c0a0a62-6ac0-4469-946d-d3c8e636b7c9_1294x468.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!HogY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c0a0a62-6ac0-4469-946d-d3c8e636b7c9_1294x468.png" width="1294" height="468" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3c0a0a62-6ac0-4469-946d-d3c8e636b7c9_1294x468.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:468,&quot;width&quot;:1294,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:52852,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/212614385?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c0a0a62-6ac0-4469-946d-d3c8e636b7c9_1294x468.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!HogY!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c0a0a62-6ac0-4469-946d-d3c8e636b7c9_1294x468.png 424w, https://substackcdn.com/image/fetch/$s_!HogY!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c0a0a62-6ac0-4469-946d-d3c8e636b7c9_1294x468.png 848w, https://substackcdn.com/image/fetch/$s_!HogY!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c0a0a62-6ac0-4469-946d-d3c8e636b7c9_1294x468.png 1272w, https://substackcdn.com/image/fetch/$s_!HogY!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3c0a0a62-6ac0-4469-946d-d3c8e636b7c9_1294x468.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Sidecar clustered </span><strong><span>20&#8211;22s</span></strong><span>. CI sat on </span><strong><span>69s</span></strong><span> every time. That stability is either a real push-and-queue signature in our harness or a coincidence at lab scale. I am reporting it, not universalizing it.</span></p><p><strong><span>p95 TTS:</span></strong><span> sidecar ~23&#8211;25s; CI ~72&#8211;99s (tails worse on push-per-task).</span></p><h3><span>LLM cost comparison explained</span></h3><p><span>The speed table is not the only controlled comparison in this harness. Token spend is measured the same way: same model </span><code>claude-sonnet-4-20250514</code><span>), same ten tasks, same agent, same fix-and-retry pattern. </span><strong><span>We changed the validation path, not the work.</span></strong></p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!diOQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F464ad48b-301e-4e6a-86ad-7e7c39c1fe96_1292x272.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!diOQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F464ad48b-301e-4e6a-86ad-7e7c39c1fe96_1292x272.png 424w, https://substackcdn.com/image/fetch/$s_!diOQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F464ad48b-301e-4e6a-86ad-7e7c39c1fe96_1292x272.png 848w, https://substackcdn.com/image/fetch/$s_!diOQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F464ad48b-301e-4e6a-86ad-7e7c39c1fe96_1292x272.png 1272w, https://substackcdn.com/image/fetch/$s_!diOQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F464ad48b-301e-4e6a-86ad-7e7c39c1fe96_1292x272.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!diOQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F464ad48b-301e-4e6a-86ad-7e7c39c1fe96_1292x272.png" width="1292" height="272" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/464ad48b-301e-4e6a-86ad-7e7c39c1fe96_1292x272.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:272,&quot;width&quot;:1292,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:32682,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/212614385?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F464ad48b-301e-4e6a-86ad-7e7c39c1fe96_1292x272.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!diOQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F464ad48b-301e-4e6a-86ad-7e7c39c1fe96_1292x272.png 424w, https://substackcdn.com/image/fetch/$s_!diOQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F464ad48b-301e-4e6a-86ad-7e7c39c1fe96_1292x272.png 848w, https://substackcdn.com/image/fetch/$s_!diOQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F464ad48b-301e-4e6a-86ad-7e7c39c1fe96_1292x272.png 1272w, https://substackcdn.com/image/fetch/$s_!diOQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F464ad48b-301e-4e6a-86ad-7e7c39c1fe96_1292x272.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a></figure></div><p><span>Both arms hit the </span><strong><span>same lint and test failures</span></strong><span> and applied the </span><strong><span>same corrections</span></strong><span>. Sidecar surfaced them via </span><code>chunk validate --remote</code><span> on the warmed snapshot (~22s median). The CI arm surfaced them after push to GitHub and a </span><strong><span>~3&#215; longer</span></strong><span> wait for the same gate jobs (~69s median). Same signal, different delivery speed.</span></p><p><span>That is why token spend landed flat. </span><strong><span>Tokens track fixes, not wait.</span></strong><span> The CI arm cost more clock time, not more tokens. Sidecars did not shrink the LLM bill here because there was nothing to shrink: identical failure set, identical agent work on both arms. On a messier repo, especially if remote CI surfaces failures the sidecar doesn&#8217;t mirror, token spend might diverge. This harness didn&#8217;t test that.</span></p><p><span>So treat </span><strong><span>~$0.90 per run</span></strong><span> as what this harness actually cost for this specific experiment, not a budget for production agent loops. Whether faster feedback trims token burn on messier work and CI failure modes (i.e. given fewer stale-context retries, less thrash, and less environment drift) is a follow-on question for a future study (coming soon). This race did not answer it.</span></p><p><strong><span>This experiment answers a narrow question cleanly: when the agent does the same fixes either way, faster feedback does not change token spend. It changes how long the agent sits idle.</span></strong></p><h3><span>Same gate jobs. Different machinery.</span></h3><p><span>We timed the same </span><strong><span>question</span></strong><span> on both arms: did </span><code>lint</code><span> and </span><code>test</code><span> pass? We did </span><strong><span>not</span></strong><span> run identical validation. Sidecar answers via </span><code>chunk validate --remote</code><span> on a warmed-up Linux snapshot. CI answers via real CircleCI jobs after you push.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!MGSu!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5298c8af-80b5-47e9-84cb-68a3851551a2_1286x744.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!MGSu!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5298c8af-80b5-47e9-84cb-68a3851551a2_1286x744.png 424w, https://substackcdn.com/image/fetch/$s_!MGSu!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5298c8af-80b5-47e9-84cb-68a3851551a2_1286x744.png 848w, https://substackcdn.com/image/fetch/$s_!MGSu!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5298c8af-80b5-47e9-84cb-68a3851551a2_1286x744.png 1272w, https://substackcdn.com/image/fetch/$s_!MGSu!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5298c8af-80b5-47e9-84cb-68a3851551a2_1286x744.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!MGSu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5298c8af-80b5-47e9-84cb-68a3851551a2_1286x744.png" width="1286" height="744" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/5298c8af-80b5-47e9-84cb-68a3851551a2_1286x744.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:744,&quot;width&quot;:1286,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:145190,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://www.confidentcommit.com/i/212614385?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5298c8af-80b5-47e9-84cb-68a3851551a2_1286x744.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!MGSu!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5298c8af-80b5-47e9-84cb-68a3851551a2_1286x744.png 424w, https://substackcdn.com/image/fetch/$s_!MGSu!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5298c8af-80b5-47e9-84cb-68a3851551a2_1286x744.png 848w, https://substackcdn.com/image/fetch/$s_!MGSu!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5298c8af-80b5-47e9-84cb-68a3851551a2_1286x744.png 1272w, https://substackcdn.com/image/fetch/$s_!MGSu!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F5298c8af-80b5-47e9-84cb-68a3851551a2_1286x744.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>Early iterations on both arms still saw </span><strong><span>lint failures</span></strong><span> while the agent fixed mistakes (about two failing lint iters per run in the rollup). The race is not &#8220;sidecar never fails.&#8221; It is </span><strong><span>how long you wait to learn</span></strong><span>.</span></p><p><span>Illustrative lab math only: ten tasks &#215; ~47s saved &#8776; </span><strong><span>~8 minutes less gate-waiting per replicate</span></strong><span> before the sidecar epilogue. Your mileage will vary.</span></p><h2><span>Takeaway</span></h2><p><strong><span>Green PRs are the outcome. Time to signal is the throttle.</span></strong></p><p><span>The </span><a href="https://loop.circleci.com/afk-builds-with-100-green-prs-chunk-sidecars-inside-the-agent-loop"><span>May experiment</span></a><span> showed sidecars help agents </span><strong><span>stop polluting branches</span></strong><span> with commits CI would reject. This one shows why that matters in wall clock terms: </span><strong><span>the coding agent is not stuck in queue while context goes cold.</span></strong></p><p><span>What held across ten runs in </span><a href="https://github.com/CircleCI-Public/chunk-cli"><span>chunk-cli</span></a><span>:</span></p><ol><li><p><strong><span>Sidecars save feedback latency</span></strong><span> on mirrored gate jobs (~3.1&#215; median TTS here).</span></p></li><li><p><strong><span>LLM spend was flat</span></strong><span> (~$4.64 sidecar vs ~$4.73 CI) in the same apples-to-apples comparison: same failures, same fixes, CI just took ~3&#215; longer to surface them. Tokens track fixes, not wait.</span></p></li><li><p><strong><span>Push-per-task CI is a predictable wait</span></strong><span> in this harness (69s median, every replicate).</span></p></li><li><p><strong><span>You still want an outer loop.</span></strong><span> Sidecar runs finished with a </span><strong><span>full ci workflow</span></strong><span> epilogue. Fast inner loop on each task, plus an authoritative pipeline confirmation at the end. Not sidecar-only cowboy coding.</span></p></li></ol><p><strong><span>Please re-read #4 above: I am not at all suggesting you delete your CI pipelines.</span></strong><span> Quite the opposite: I am saying that if your agent loop is </span><strong><span>push &#8594; wait &#8594; fix &#8594; push again</span></strong><span>, you are paying for that wait </span><strong><span>every iteration</span></strong><span>.</span></p><p><span>Sidecars moved the same exact question to a local, warmed-up Linux snapshot...</span></p><p><strong><span>QUESTION:</span></strong><span> </span><em><span>Are lint and test happy on this repo right now?</span></em></p><p><span>...and the same answers came with a much tighter feedback loop. </span><strong><span>3.1&#215; faster.</span></strong></p><h2><span>What&#8217;s Next</span></h2><ul><li><p><span>Combine this harness with </span><a href="https://loop.circleci.com/hardening-ralphci-loops-for-open-source-after-the-february-2026-study"><span>RalphCI</span></a><span> orchestration (</span><a href="https://loop.circleci.com/afk-builds-with-100-green-prs-chunk-sidecars-inside-the-agent-loop"><span>Snake outcome study</span></a><span> + TTS race in one stack).</span></p></li><li><p><span>Heavier repos, flaky tests, parallel agents. Find where Chunk sidecars and snapshot sync stops winning.</span></p></li><li><p><strong><span>Agent economics at scale</span></strong><span>: cost per green iteration (LLM + validation path), not just TTS. A proper cohort study on heavier repos, not extrapolated from this ten-run table alone.</span></p></li><li><p><span>Demo recordings on various Chunk sidecar and snapshot sync setups. </span><a href="https://www.youtube.com/watch?v=P99Bk8bQRsA&amp;list=PL9GgS3TcDh8xdRpucbu7Y7dq6lTlx7Pju"><span>Stay tuned!</span></a></p></li></ul><p><strong><span>Artifacts:</span></strong></p><ul><li><p><a href="https://github.com/CircleCI-Public/chunk-cli/blob/sidecar-race-05-27-2026/experiments/sidecar-race/FINDINGS.md"><span>FINDINGS.md</span></a></p></li><li><p><a href="https://github.com/CircleCI-Public/chunk-cli/pull/370"><span>Experiment PR #370</span></a></p></li></ul><p></p><h4></h4>]]></content:encoded></item></channel></rss>