Internally, CircleCI has been focused on improving our platform for AI agents. As part of that focus, we’ve been running different experiments evaluating the performance of AI agents around a specific problem: fixing a broken pipeline (one of, if not the most common issue of our customers).
After ~45 experiments we reached a plateau: we were getting a consistent 85% +/- 3% pipeline fix rate in different runs, no matter how the agent fetched the failed pipeline context. As a result, we wanted to further examine the details of each run, asking ourselves the following questions:
How many samples does the agent always fix?
How many samples does the agent never fix?
How many samples would switch fix-result run over run?
The results were a surprise: 60% were always fixable, 15% were never fixed, and 25% would be fixed in some runs but not in others. We named the last share flip rate: the percentage of tasks whose outcome changes across repeated runs on the same input. A high flip rate means any single run is unreliable: differences between experiments may be noise rather than a real effect.
The next question became: Why? What changed across runs that caused this variation of outcomes?
We looked at the different fixes the agent did to these swapping samples and the answer was more simple than we predicted: This variation was a result of agentic non-determinism; the agent would sometimes come up with the correct solution, but sometimes it would miss certain aspects of the fix, or try something completely different. Note that this is distinct from hallucination since agents were not making up a fix; they were looking at different paths that seemed valid and landed in different places.
With this in mind, we moved on to creating a solution. How can we make more consistent fixes? Which tools can help with this? As we started researching what other people were doing, a trend emerged: Helping an agent better manage its memory/context reduces the unexpected behavior and creates better quality outputs.
For more information on the research we were consulting to inform our work, check out code rabbit memory and Mendral.
Hypothesis
The agent fails to be consistent across runs, so giving the agent an external memory should help reduce this inconsistent work.
Before stepping into developing a full context graph or anything too complex, I decided to try the least complicated implementation: a simple note-taking tool. I anticipated that this would give the agent a way to annotate its most important findings and read them at will, thus producing better and more consistent output.
Setup
In designing the project’s structure, we wanted a way to prove that what we added actually had an impact. As I’ve mentioned in a past article, I’m a fan of the scientific method and a big believer in repeatable structure. So, we designed this as a formal experiment with a baseline and a target to measure against. Also, due to the non-deterministic behavior of agents, we made sure to run the same experiment multiple times to validate that our results maintained across runs.
Two tools were added to the agent:
A tool to save a note with a given key, similar to a dictionary.
A function to return all the stored notes.
We tried three different setups:
Basic: Give the tools an instruction to “save each finding, read before finalizing”.
Formal read protocol: Tools and instruction to read before diagnosis and before validation (mandatory).
Auto-surface: When saving a note, it would return the full accumulated notes on every call, reaching context automatically.
Across the 25% fix rate samples, we tested this against Claude Code with Sonnet 5, running each setup three times to measure actual flip rate. Each run started fresh, so no context is shared across runs. To measure success, we ran the real pipeline to capture outcome, and we also used an Agent-As-Judge (Claude Opus 4.6) to evaluate the quality of the fix, focusing on if it was “mergeable or not” (we call this “merge quality”).
Results
The three mechanisms reduce flip-rate on a certain level, producing more consistently passing pipelines across attempts, but fix rate and merge quality had a different behavior across setups.
Basic setup dropped on flip rate, but hurt on quality. Looking into the behavior of the agent, it was writing a lot, but it almost never read what it wrote. This affected the solution’s quality, dropping from 72% from the control experiment, to 60%.
The second setup was the other way around, it achieved the biggest reduction in flip-rate (from 50% to 25%) and quality remained the same. Adoption from the agent was visible too, with all the sessions reading the notes as instructed.
Finally, the third setup also decreased the flip rate to 30% and improved quality by a small margin, proving again that giving the agent a place to better manage its memory/context helps to reduce the uncertainty of its behavior.
Since the goal was also to have the smallest impact possible in token usage, we tested the second setup with the full dataset, three times, to measure actual impact on reducing the flip-rate without having regressions. The results can be found in the table below.
The benefits are clear: reducing the flip rate by ⅓ without dropping fix rate nor quality, costing less than 10% more tokens per sample, and taking less than 5% more time.
Sidebar: In a separate experiment, we found that testing with Claude Opus 5 showed a fix rate improvement of +5% and quality improvement of +7%, at 36% higher cost, which is worth knowing before switching models.
TL;DR
It is clear that nowadays we cannot rely only on the agent’s context to get the work done.
As problems grow, external memory tools do improve on achieving more consistent results. These experiments back up that trend: reducing the variability of results with only a 7% increase on token consumption is a really good deal, rather than using a bigger (more expensive) model with a bigger context window.
Our findings here mean we can now consider other memory tools in our day-to-day agent usage.




