lyr.ai

Agent reliability · Method note 03

Why agent CI can't be treated like deterministic tests

Code for this post: lyr-ai/agentseism

Continuous integration rests on one assumption so basic we rarely say it out loud: run the same code on the same input and you get the same result. A test that passed on main and fails on your branch is evidence against your branch. That assumption is what makes a red check mean something.

AI agents break it.

Run an agent twice on the same task, with the same code, model and prompt, and it can read different files, call tools in a different order and end somewhere else. Sometimes it succeeds both times, sometimes once. So when a pull request touches the agent and the eval score moves, what should the check believe?

A check that failed for the right reason

Here is a pull request that looks like a harmless refactor. It changes one line in a checkout handler, and the field it now reads, amount_cents, does not exist.

A GitHub pull request, 'checkout: use amount_cents for the receipt amount'. The github-actions bot has posted an AgentSeism report headed 'REGRESSION — do not merge without review'. Its scorecard has two rows. Task success (broad): baseline 0.93, PR 0.68, change -0.25 with interval -0.52 to -0.04, decision PASS. Task success (capability): 7 of 7 tasks monitored, 1 collapsed, checkout 8/8 to 0/8, decision REGRESSION. The evidence list shows checkout fired with p=0.0001 and a non-blocking warning on return-label, 7/8 to 3/8.
Figure 1. A public pull request in an example repository. The agent there is simulated: real handler code, fixed per-task success rates, and outcomes that vary from run to run. The CI check around it is the real one.

The check ran each of seven tasks eight times on main and eight times on the branch. Overall success fell from 0.93 to 0.68, and two things happened that ordinary CI has no vocabulary for:

A third detail is easy to miss: return-label went from 7/8 to 3/8 and was flagged as a warning, not a failure. Its success rate didn’t change in this PR, so that drop is pure run-to-run wobble, and a warning is the right amount of alarm for it.

A single pass/fail test cannot express any of that. Neither can a single eval score.

One run is a sample, not a replay

A deterministic test is a replay: run it again and you learn nothing new. An agent run is a sample from a distribution of possible runs. That changes what a comparison between two branches can tell you, in three ways.

Three cards. Normal noise: task success 91% to 86%, PASS, ordinary run-to-run wobble. Capability collapse: 91% to 77%, REGRESSION, checkout 8/8 to 0/8. Not enough evidence: 100% to 83%, NEED EVIDENCE, 3 tasks with 2 runs per side.
Figure 2. The three situations a stochastic check has to tell apart. These come from seism demo, which feeds fixed, stated outcomes through the real decision engine. They illustrate the cases and are not evidence.

1. Scores move when nothing changed

On a real coding agent (mini-swe-agent with Claude Haiku 4.5, on five SWE-bench tasks run five times each), we ran an unchanged candidate against its own baseline. Success went from 0.92 to 0.88. One task went from 3/5 to 2/5 with no code change at all.

A deterministic mindset reads that as “the PR made it worse”. It didn’t. There was no PR. A check that blocks on this trains developers to click “re-run” until it goes green, and after that nobody trusts it.

2. Averages hide breakage

The opposite failure is quieter. In the pull request above, the average fell 25 points, but not evenly: most of it was one task collapsing completely. Averaged over the suite, a total failure of one capability looks like a moderate, uncertain drop.

A measurement plate of seven capabilities, each with eight runs on main and eight on the pull request. Standing marks are successful runs; dots on the floor are failures. Six capabilities barely change. Checkout goes from eight standing marks on main to eight red dots on the pull request, 8/8 to 0/8. Beneath, a seismic trace shows small tremors across the suite and then a sharp fall at checkout. The average fell 14 points; checkout fell 100.
Figure 3. The average fell 14 points. One capability disappeared. ▶ Explore what happened: an interactive version of the seism demo collapse scenario, where every mark is one run.

This isn’t hypothetical for us. The first version of our own check measured only the suite-wide average, and on a real agent it passed a change that cut success from 0.92 to 0.52, with two tasks collapsing. The statistics were computed correctly. They were answering the wrong question. That failure deserves its own post.

3. Sometimes the honest answer is “not enough evidence”

Three tasks, run twice each, going from 100% to 83% is not a regression and not a pass. It is too little data. A binary pass/fail setup has no honest way to represent that, so it gets rounded to whichever side of the threshold it falls on.

Why not just run your eval more times?

Running more repetitions is the obvious fix, and it helps, but it doesn’t answer the question on its own. More numbers still leave four decisions open:

Repeating the eval gives you the evidence. Something still has to turn that evidence into a merge decision.

Turning repeated runs into a merge decision

That is what AgentSeism does. You give it a command that runs your agent on a task and one that checks the result. It runs your base branch and your PR several times each, keeps each task’s results together, and returns one of four verdicts:

Verdict Meaning
PASS Neither gate found enough evidence of a material regression.
REGRESSION The suite as a whole is confidently worse by at least your threshold, or a task that reliably worked has collapsed.
INSUFFICIENT EVIDENCE Too little data to decide. It blocks the merge, but it is not reported as a regression.
INCOMPARABLE The model, runtime or dependencies differ between the two sides, so nothing is compared.

The two gates behind REGRESSION answer different questions, which is why they are separate: did the suite get broadly worse? and did something that used to work stop working? The pull request above is the case where the answers differ.

One correction to the previous note in this series. It ended on “Comparability first. Outcomes decide. Traces explain.” The first two became the product. The third, localising a regression to a stage of the agent’s trajectory, did not, and AgentSeism does not claim it.

How much should you trust it?

We froze the decision rule before testing it, then ran it on seven SWE-bench tasks it had never seen, with a pre-registered protocol: 224 real agent runs, 0 invalid, $38.49 in API cost. It passed an unchanged agent, and it caught both degradations that had been predicted in advance, a capability collapse and a broad collapse.

That evidence is narrow, and it is worth saying how:

The full protocol and results, including what the study did not establish, are in the repository.

Try it

If your agent’s eval score has ever moved on a PR and you couldn’t tell whether to believe it:

git clone https://github.com/lyr-ai/agentseism.git && cd agentseism
python3.11 -m venv .venv && source .venv/bin/activate
pip install -e .
seism demo

It takes about a second, with no API key and no Docker. Then look at the two public pull requests in agentseism-example: one docs-only change that passes, and the one-line bug above.

Or run your own eval eight times on an unchanged branch, and see how stable it really is.

If this is a problem you have too, a ⭐ on the repo tells me it’s worth continuing. An issue telling me where it breaks on your agent is worth even more.


Written while building AgentSeism. New posts by RSS.