I handed a coding agent seven tasks and a classic Snake game spec, then walked away from the keyboard to let it cook. AFK, start to finish. Every task iteration that cleared inner-loop checks got pushed, and the outer-loop CI pipeline ran on every one of those pushes.
When a pipeline came back red, the local fix loop got triggered: read the failure, patch it, push again.
That loop is where an AFK run burns dollars and wall-clock nobody budgeted for. And a fat slice of that bill is nothing more than the CI pipeline failure context we shovel into the model when the agent asks for it.
So I ran the build again and changed exactly one thing: what the fix agent reads when CI goes red. Push cadence held still. Same model, same task list, same expectation (a merge-ready green PR at the end of the run).
One arm was given the entire raw CI pipeline failure output. The other got a compact, structured report from the following CircleCI CLI command: circleci run get --failure-report.
Full pipeline failure output: ~800,898 characters.
--failure-report: ~19,936.
About 40x less stuff stuffed into the prompt.
Across both arms, fix loops stayed at 3, and pipelines stayed at 8.
That is the finding. Compact, structured failure context makes a retry cheaper. It does not, by itself, make retries disappear.
Nothing here prevents the red. It makes the red cheaper to read.
Hypothesis
If the CI fix loop reads a compact, structured failure report instead of the whole log dump, CI pipeline failure diagnosis input tokens should drop. Retry count should not.
Same red pipelines fed to the fix agent. Thinner diet.
Setup
Same Snake build both arms. Seven tasks. AFK. Push after every task that cleared inner-loop checks.
The variable was failure shape, not how often we triggered outer-loop CI.
per-task-push (full log): when a pipeline went red, the fix agent fetched the entire CI pipeline failure output.
per-task-push
--failure-report): same cadence. Failures shaped with the CircleCI CLI commandcircleci run get --failure-report. Compact. Structured. Sized for an agent.
Everything else that could wander, I tried to hold still: model, task list, merge-ready green PR at the end of the run.
Results
Both arms finished green. Both still took the outer pipeline eight times. Both still ran the fix loop three times.
So why care about fewer failure-log characters if cost, wall-clock, pipelines, and fix loops look basically the same?
Because the totals are misleading.
Most of that ~$15 bill is the agent writing the game, not reading CI failures. That build cost jumps around run to run (call it $10 to $17). So even when we cut failure-log input from ~800k characters to ~20k, the total check can only drop about $1 to $2. Wall-clock barely moves either. Reading a thinner failure report might save milliseconds or seconds, not minutes.
The savings are still real on every fix. Every time the agent has to diagnose a red pipeline, you stuff ~40x less into the prompt. That is input tokens you do not buy. On this pair of runs the whole-run LLM bill moved from $16.75 to $14.27. Do not hang a strategy on that exact delta. Codegen variance is loud. The quiet, repeatable tell is the input diet: ~800k → ~20k characters of failure context, with zero change in how often we came back for another CI pipeline run.
Extremely detailed feedback on every attempt. Same number of attempts. Cheaper to read each time.
TL;DR
--failure-report is graded comments, not the whole exam booklet: failed steps, condensed output, enough to fix without hauling the universe into the context window.
So treat it as hygiene. Use the flag. Always. Pull it from the CircleCI CLI, pipe it into the fix loop, point the agent at it, and stop letting that loop read raw logs. It belongs in any decent agent harness. For the longer cut on tokens and fix time versus raw logs, see the --failure-report write-up.
Cheaper retries are good. Zero retries are better.
If the goal is a green PR in one take, this flag will not get you there. Showing the agent what CI will check before it writes product code will. That is a different experiment, and it is the one that drove fix loops from 3 to 0.



