<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.10.0">Jekyll</generator><link href="https://lyr-ai.github.io/feed.xml" rel="self" type="application/atom+xml" /><link href="https://lyr-ai.github.io/" rel="alternate" type="text/html" /><updated>2026-09-28T11:34:01-07:00</updated><id>https://lyr-ai.github.io/feed.xml</id><title type="html">lyr.ai</title><subtitle>Research notes on why AI agents behave inconsistently, and what can be measured about it. Written alongside the code: AgentSeism, TypedMem, LYR.</subtitle><author><name>Ruxi Zhang</name></author><entry><title type="html">When your AI supplier becomes your competitor</title><link href="https://lyr-ai.github.io/when-your-ai-supplier-becomes-your-competitor/" rel="alternate" type="text/html" title="When your AI supplier becomes your competitor" /><published>2026-09-27T20:00:00-07:00</published><updated>2026-09-27T20:00:00-07:00</updated><id>https://lyr-ai.github.io/when-your-ai-supplier-becomes-your-competitor</id><content type="html" xml:base="https://lyr-ai.github.io/when-your-ai-supplier-becomes-your-competitor/"><![CDATA[<p>Harvey sells AI to lawyers. It’s one of the most valuable application companies built on
top of frontier models: in September 2026 it raised $550M at a $15.5B valuation, and its
co-founder said it had crossed $400M in annual recurring revenue.</p>

<p>This year, the companies whose models Harvey builds on started selling legal AI too.
Anthropic released its first legal plugin at the end of January and expanded it into Claude
for Legal in May. In September OpenAI released Astra for Law, a version of its newest model configured for
legal work.</p>

<p>So Harvey’s suppliers are becoming its competitors. Yet Harvey’s reported ARR continued to
rise sharply. And in the same period, rather than replacing the labs, Harvey has been
investing in more of the layer between their models and legal work: benchmarks,
post-trained models, workflows, and people inside law firms.</p>

<p>This piece looks at how Harvey got here, where it sits, and what its valuation assumes.
The question it ends on is whether owning more of that layer is enough.</p>

<h2 id="how-harvey-got-here">How Harvey got here</h2>

<figure>
  <img src="/assets/img/harvey-timeline.png" alt="Timeline from November 2022 to September 2026 in four lanes. Product: assistant on GPT-4, Vault, agents, multi-model, Shared Spaces; below the app layer, a custom model with OpenAI in 2024, then in 2026 an open benchmark (LAB), post-training projects and Tenet (research preview). The labs in legal: nothing before 2026, then a Claude legal plugin in January, Claude for Legal in May, Astra for Law in September. Reported ARR: more than $50M in February 2025, more than $100M in August 2025, $190M in January 2026, more than $400M in September 2026. Valuation: $0.7B, $1.5B, $3B, $5B, $8B, $11B, $15.5B." loading="lazy" />
  <figcaption><b>Figure 1.</b> In 2026, three things moved at once: the labs entered legal, Harvey started building below the application layer, and its revenue grew fastest.</figcaption>
</figure>

<p>For its first three years, Harvey’s product remained primarily an application layer built on
other companies’ foundation models.
It started as a legal assistant on GPT-4, added a document store (Vault), then agents and
workflows. In 2025 it went from one model provider to several, adding Anthropic’s and
Google’s models alongside OpenAI’s.</p>

<p>Its business grew quickly. The company reported more than $50M in ARR in February 2025,
more than $100M that August, $190M in January 2026, and more than $400M in September 2026.
Each new round was priced higher: $3B, $5B, $8B, $11B and now $15.5B, all within about 19
months.</p>

<p>2026 is where the lanes line up. The labs had no legal products before this year; by
September both Anthropic and OpenAI had one. In the same months Harvey did things an
application company usually doesn’t:</p>

<ul>
  <li>it open-sourced a legal benchmark (LAB) in May;</li>
  <li>it ran post-training projects on open-weight models with partners from May to August;</li>
  <li>in August it released Tenet, “our first post-trained open-weight model”, as a research
preview.</li>
</ul>

<p>The timing is striking, but it doesn’t show Harvey was reacting to the labs. Harvey also
tends to disclose its ARR around funding rounds, so the business lane is partly what the
company chose to announce.</p>

<h2 id="where-harvey-sits">Where Harvey sits</h2>

<figure>
  <img src="/assets/img/harvey-supplier-and-competitor.png" alt="Layered map. Top layer, models: OpenAI and Anthropic, open-weight bases (Kimi, GLM, Qwen), RELX/LexisNexis legal content. Middle layer, legal AI products: Harvey, Legora, Thomson Reuters. Bottom layer: law firms and legal teams, and in-house builds such as Freshfields on Claude. The labs supply models to Harvey, host Harvey as a plugin in ChatGPT, and sell legal products (Claude for Legal, Astra for Law) directly to law firms. Harvey sells seats and deployment to law firms. Harvey's LAB benchmark and Tenet research preview are drawn dashed. Legora and Thomson Reuters compete for the same customers; RELX licenses content to Harvey." loading="lazy" />
  <figcaption><b>Figure 2.</b> The same labs play three roles around Harvey. Other competitors enter from different layers.</figcaption>
</figure>

<p>The unusual part isn’t that frontier labs compete with Harvey. It’s that the same companies
now stand in three places around it at once.</p>

<p><strong>They supply its models.</strong> Harvey’s own benchmark writing describes a version of its
Assistant as “built primarily on GPT-5”. OpenAI says Astra for Law will be “available soon
to API customers, including Harvey and Legora”. Harvey’s newest capabilities will partly
come from the same company that competes with it.</p>

<p><strong>They host it.</strong> Alongside Astra, OpenAI launched legal plugins for ChatGPT, “26 from
vendors such as Thomson Reuters, Harvey, Legora and iManage”. Harvey is now also something
you can reach from inside a lab’s product.</p>

<p><strong>They sell to its customers.</strong> Anthropic’s Claude for Legal comes with legal plugins and
connectors for specific areas of law. Astra for Law is offered first to selected large
firms. Freshfields has deployed Claude across the firm and is co-building legal workflows
with Anthropic.</p>

<p>Around that centre, the pressure comes from other directions:</p>

<ul>
  <li><strong>Legora</strong>, another AI-native legal startup, is valued at $5.55B after its spring round,
with more than $200M in ARR reported this month.</li>
  <li><strong>Thomson Reuters</strong> (CoCounsel, Westlaw) and <strong>RELX/LexisNexis</strong> own the legal content
and the distribution lawyers already pay for. RELX’s LexisNexis also licenses content into
Harvey.</li>
  <li><strong>Large firms can build their own.</strong> Freshfields is working directly with Anthropic, and
Kirkland &amp; Ellis has reportedly committed to its own platform with Palantir.</li>
</ul>

<h2 id="what-harvey-actually-owns">What Harvey actually owns</h2>

<p>Judging by what Harvey sells and where it is hiring, it still looks primarily like an
application and deployment company:</p>

<ul>
  <li><strong>Customers.</strong> The co-founder says Harvey has more than 3,000 customers, including 80% of
the top 100 US law firms, 20% of the Fortune 500 and half of the Fortune 10. How many of
those are firm-wide deployments isn’t disclosed.</li>
  <li><strong>Workflow product.</strong> Assistant, Vault, agents and collaboration features. That’s the
part lawyers use every day.</li>
  <li><strong>People inside firms.</strong> Of 304 open roles on its job board, 33 are for legal engineers,
people who deploy Harvey into a firm’s work. About six or seven are for research or model
training.</li>
</ul>

<p>Below the app, it owns less than the headlines suggest. It hasn’t pre-trained a model. Its
post-trained models start from other companies’ open weights (Moonshot’s Kimi, Zhipu’s GLM,
Alibaba’s Qwen) and are trained with partners. Tenet is a research preview, and nothing
public shows it serving production traffic. Its benchmark is open, and its gains are
measured on that benchmark.</p>

<p>Harvey does give reasons for going down the stack. On its own benchmark, “reaching the top
of the closed-source leaderboard runs to roughly $50 per task and over 20 minutes of
latency”, and frontier models complete “less than 10% of tasks end-to-end”. Open-weight
models “can be hosted within a firm’s own secure cloud environment”. After the latest round,
its co-founder said the company would invest “heavily in both its harness and its own model
training”.</p>

<p>So the honest description today is an application company with a growing bet below the
application layer. The bet is not yet a business.</p>

<h2 id="what-investors-are-pricing-in">What investors are pricing in</h2>

<figure>
  <img src="/assets/img/harvey-valuation-multiples.png" alt="Dot plot on a log scale of valuation divided by revenue. Harvey by round: at most 60 times in February 2025, 50 to 67 times in June 2025, about 42 times in December 2025, about 58 times in March 2026 (denominator from an earlier date), at most 39 times in September 2026. Legora: about 55 times in spring 2026, about 42 times in reported September talks. Thomson Reuters and RELX: 5.4 to 5.8 times EV/Sales for the whole company. Not directly comparable: shown for the scale of expectations." loading="lazy" />
  <figcaption><b>Figure 3.</b> Private legal-AI companies are valued at roughly 40–60 times reported revenue; the incumbents at about 5–6 times sales. The two aren't directly comparable; the gap shows how different the expectations are.</figcaption>
</figure>

<p>At $15.5B and more than $400M in ARR, Harvey is valued at no more than about 39 times its
recurring revenue. That’s lower than at its earlier rounds, because reported ARR grew faster
than the valuation, but it’s still far from the incumbents. Thomson Reuters and RELX trade at
roughly 5–6 times their sales. Legora sits in the same range as Harvey.</p>

<p>These numbers measure different things (a private round’s valuation against reported ARR,
against a public company’s enterprise value over a year of total revenue), so the
comparison isn’t precise. What it shows is scale: the multiples imply expectations very
different from those attached to mature legal-information businesses.</p>

<p>That leads to the useful question. It isn’t “is Harvey worth $15.5B?” but <strong>what has to
become true for Harvey to grow into these expectations?</strong></p>

<h2 id="three-paths">Three paths</h2>

<p>The evidence supports three plausible directions. None of them is decided, because the
numbers that would decide them (margins, retention, and whether Harvey’s own models carry
real traffic) aren’t public.</p>

<p><strong>1. A vertical intelligence layer.</strong> Harvey’s own benchmark, models and data become the
reason firms stay.</p>
<ul>
  <li><em>Must become true:</em> its own models serve a meaningful share of production work, and that
shows up as better margins or prices; firms pay for firm-specific models.</li>
  <li><em>Breaks it:</em> the next frontier generation beats post-trained open models at similar cost,
or firms post-train their own.</li>
</ul>

<p><strong>2. The deployment platform on other people’s models.</strong> Harvey stays model-agnostic and
wins on workflow, integration and the people it puts inside firms.</p>
<ul>
  <li><em>Must become true:</em> retention and expansion hold as the labs sell direct, and deployment
depth grows faster than the labs’ own legal offerings.</li>
  <li><em>Breaks it:</em> firms standardise on a lab plus an internal platform, as Freshfields is doing
with Anthropic, or usage pricing compresses revenue per seat.</li>
</ul>

<p><strong>3. Beyond law.</strong> Harvey becomes a platform for professional services more broadly. It has
bought an asset-management platform and says it works with 125+ asset managers.</p>
<ul>
  <li><em>Must become true:</em> non-legal revenue becomes a material, disclosed share.</li>
  <li><em>Breaks it:</em> incumbents or the labs own those workflows first. The evidence here is
mostly announcements.</li>
</ul>

<h2 id="what-limits-every-path">What limits every path</h2>

<ul>
  <li><strong>The technology isn’t done.</strong> On Harvey’s own benchmark, frontier models complete less
than 10% of long tasks end to end, and Tenet improves on its base model without getting
close.</li>
  <li><strong>The labs move up.</strong> Their legal products, and their models inside competitors’
products, reduce what an application layer adds.</li>
  <li><strong>The incumbents move down.</strong> Thomson Reuters and RELX have the content and the contracts
lawyers already rely on.</li>
  <li><strong>Customers can build.</strong> The largest firms have the money and, increasingly, the partners
to do it themselves.</li>
</ul>

<h2 id="the-open-question">The open question</h2>

<p>We know Harvey has built distribution and revenue quickly. We know it’s experimenting
below the application layer. We don’t yet know whether owning more of the model stack is
necessary for its business, or even valuable to it.</p>

<p>The open question isn’t whether Harvey can build a legal model. It’s whether owning more of
the intelligence layer makes the company harder to replace than simply owning the customer
relationship.</p>

<h2 id="what-we-know">What we know</h2>

<table>
  <thead>
    <tr>
      <th>Evidence level</th>
      <th>Claim</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Observed</td>
      <td>The labs launched legal products in 2026; OpenAI’s legal launch includes Harvey both as an API customer and as a ChatGPT plugin; Harvey released an open benchmark and a post-trained model in research preview; far more of its open roles are for legal engineers than for model research</td>
    </tr>
    <tr>
      <td>Reported</td>
      <td>$15.5B valuation; more than $400M ARR and 3,000+ customers (company); earlier ARR points and round valuations; competitors’ valuations and ARR</td>
    </tr>
    <tr>
      <td>Inferred</td>
      <td>Harvey is hedging below the application layer rather than replacing its suppliers; the valuation prices in growth very unlike an incumbent’s</td>
    </tr>
    <tr>
      <td>Unknown</td>
      <td>Margins, retention, how much work runs on Harvey’s own models, what “customer” counts, contract terms with the labs and LexisNexis</td>
    </tr>
  </tbody>
</table>

<hr />

<p><em>Sources: Anthropic’s public knowledge-work-plugins repository (the legal plugin is in its
first commit, 2026-01-29); Harvey’s blog (the Tenet research preview; post-training with Baseten; BigLaw
Bench Arena); LawSites on the September 2026 round and on Astra for Law; TechCrunch and
PointBlank on Claude for Legal; Freshfields on its Anthropic partnership; Harvey’s job
board (Ashby, read 2026-09-28); public-market multiples as of 2026-09-27. Revenue figures
are company-reported unless marked; where sources disagree, the research traced each figure
to its date and definition rather than picking one.</em></p>]]></content><author><name>Ruxi Zhang</name></author><summary type="html"><![CDATA[Harvey is growing fast between frontier labs moving up into legal software and legal incumbents moving down. Its answer so far isn't to replace its model suppliers: it's to own more of the layer between models and legal work. Is that enough?]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://lyr-ai.github.io/assets/img/harvey-supplier-and-competitor.png" /><media:content medium="image" url="https://lyr-ai.github.io/assets/img/harvey-supplier-and-competitor.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">The model isn’t the system</title><link href="https://lyr-ai.github.io/the-model-isnt-the-system/" rel="alternate" type="text/html" title="The model isn’t the system" /><published>2026-09-27T18:30:00-07:00</published><updated>2026-09-27T18:30:00-07:00</updated><id>https://lyr-ai.github.io/the-model-isnt-the-system</id><content type="html" xml:base="https://lyr-ai.github.io/the-model-isnt-the-system/"><![CDATA[<p>We keep benchmarking models. Often, we’re benchmarking systems.</p>

<p>When a leaderboard says a model solves 62.9% of tasks, it’s easy to read that as the
model’s ability. But an AI coding agent is a model inside a <strong>harness</strong>: the tools it
can call, its prompts, how it manages context, the loop it runs in, the checks around
it. The score belongs to the pair.</p>

<p>How much does the harness matter? One benchmark ran the same models in different
harnesses, and the answer is: enough to flip the result.</p>

<p>On Terminal-Bench 2.0, <strong>switching from a minimal harness to each model’s vendor harness
changed scores by as much as +14 points for one model and −13 for another.</strong> And the
direction depended on which harness it was.</p>

<figure>
  <img src="/assets/img/harness-same-model-two-harnesses.png" alt="Dumbbell chart from Terminal-Bench 2.0. Each model is shown in a minimal harness (Terminus 2) and in its vendor's own harness. With Codex CLI, GPT-5 goes from 35.2% to 49.6% (+14.4), GPT-5.2 from 54.0% to 62.9% (+8.9), GPT-5-Mini from 24.0% to 31.9% (+7.9). With Claude Code, Opus 4.5 goes from 57.8% to 52.1% (−5.7), Opus 4.1 from 38.0% to 34.8% (−3.2), Sonnet 4.5 from 42.8% to 40.1% (−2.7), Haiku 4.5 from 28.3% to 27.5% (−0.8); these are within noise. With Gemini CLI, Gemini 2.5 Pro goes from 32.6% to 19.6% (−13.0), Gemini 2.5 Flash from 16.9% to 15.4% (−1.5)." loading="lazy" />
  <figcaption><b>Figure 1.</b> Every model that Terminal-Bench 2.0 ran both in a minimal harness and in its vendor's own. Faded differences are within the benchmark's reported uncertainty.</figcaption>
</figure>

<h2 id="what-the-table-shows">What the table shows</h2>

<p>The benchmark’s authors ran each model-and-harness pair at least five times and report
the uncertainty. Nine models were run both in <strong>Terminus 2</strong>, a deliberately minimal agent,
and in their vendor’s own harness:</p>

<ul>
  <li><strong>Codex CLI</strong> helped every OpenAI model: +14.4 points for GPT-5, +8.9 for GPT-5.2,
+7.9 for GPT-5-Mini.</li>
  <li><strong>Claude Code</strong> came out slightly below the minimal harness for every Claude model,
from −0.8 to −5.7 points. Each of those gaps on its own is within the reported
uncertainty; it’s the consistent direction that stands out.</li>
  <li><strong>Gemini CLI</strong> cost Gemini 2.5 Pro 13 points, and Gemini 2.5 Flash 1.5.</li>
</ul>

<p>The authors draw their own conclusion from it: “model selection is usually more
important than agent scaffold.” That’s true on average. But for a single model, the
scaffold was worth anything from −13 to +14 points.</p>

<h2 id="what-it-doesnt-show">What it doesn’t show</h2>

<p>It isn’t a ranking of harness quality. <strong>Terminus 2 was built by the benchmark’s own
authors</strong>, for this benchmark, so it plays at home. Claude Code trailing it here doesn’t
mean Claude Code is badly built, and Codex CLI’s gains here don’t transfer automatically
to your tasks.</p>

<p>It’s also one benchmark, made of terminal tasks.</p>

<p>What it does show is narrower and more useful: <strong>a harness has no fixed value on its
own.</strong> Its effect depends on the model inside it, and can change sign.</p>

<h2 id="so-when-the-model-changes-re-test-the-harness">So when the model changes, re-test the harness</h2>

<p>If the value of a harness depends on the pairing, then upgrading the model breaks your
knowledge of that pairing. The harness that helped last month’s model may be neutral, or
in the way, for this month’s.</p>

<p>The practical rule is simple: when you switch or upgrade the model, re-run the
comparison. Try your full harness against a simpler one. Remove components one at a
time. Keep what still helps.</p>

<p>And when you read a leaderboard, ask what system the number came from.</p>

<h3 id="an-ablation-not-a-comparison">An ablation, not a comparison</h3>

<p>Comparing model A in harness X with model B in harness Y tells you which pair is better,
not why. To learn what your harness is worth, hold the model fixed and take the harness
apart:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>model fixed, same tasks, several runs each

full harness                 → score, cost, time
minimal baseline             → score, cost, time
full − planning step         → ?
full − retry loop            → ?
full − custom tools          → ?
full − context rules         → ?
full − checks                → ?
</code></pre></div></div>

<p>Two things from the evidence above are worth building in:</p>

<ul>
  <li><strong>Run each configuration several times.</strong> Terminal-Bench ran every pair at least five
times and still reports roughly ±3 points. A difference smaller than that is noise.</li>
  <li><strong>Record cost and time, not just the score.</strong> A component that no longer changes
accuracy may still pay for itself by saving tokens, as the planning study found.</li>
</ul>

<p>Keep a component only if taking it out makes things measurably worse. Then run the same
table again when the model changes.</p>

<h2 id="harnesses-also-age-a-separate-weaker-signal">Harnesses also age (a separate, weaker signal)</h2>

<p>There are other signs that particular scaffolding loses its value as models improve,
though none of them come from the table above, which is a snapshot, not a history:</p>

<ul>
  <li>A study of agent design choices found that planning “shifts from an accuracy scaffold
for weaker models to a cost saver for stronger models”.</li>
  <li>METR compared Claude Code with a simple scaffold for Opus 4.5. Claude Code came out
ahead in 50.7% of their bootstrap samples: a coin flip.</li>
  <li>Anthropic, describing how it builds harnesses for long-running work, puts it this way:
“every component in a harness encodes an assumption about what the model can’t do on
its own, and those assumptions are worth stress testing.”</li>
</ul>

<p>That fits the idea that a harness moves with the model rather than simply growing. But
it’s supported by a handful of reports, not established.</p>

<h2 id="what-might-last-a-hypothesis">What might last: a hypothesis</h2>

<p>If some scaffolding goes stale, what doesn’t? Practitioners and vendors keep pointing at
<strong>verification</strong>: tests, checks and other feedback the agent can run on its own work. It’s
plausible that this ages more slowly than prompts and workflow tricks. On the evidence I
found, though, it’s a hypothesis. Nobody I found has measured it across model
generations.</p>

<h2 id="summary">Summary</h2>

<table>
  <thead>
    <tr>
      <th>How strong</th>
      <th>Claim</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Measured (one benchmark)</td>
      <td>The same model can get better or worse depending on the harness, and the direction varies</td>
    </tr>
    <tr>
      <td>Supported by several reports</td>
      <td>Some scaffolding loses value as models improve</td>
    </tr>
    <tr>
      <td>Hypothesis</td>
      <td>Verification may last longer than other parts of the harness</td>
    </tr>
  </tbody>
</table>

<p>The model isn’t the system. When you change one, re-test the other.</p>

<hr />

<p><em>Sources: Terminal-Bench 2.0 (arxiv.org/abs/2601.11868, appendix Table 2); METR,
“Measuring time horizon using Claude Code and Codex” (February 2026); the empirical study
of agent design choices (arxiv.org/abs/2609.20804); Anthropic, “Harness design for
long-running application development”. Each quote and number was read on the original page.</em></p>]]></content><author><name>Ruxi Zhang</name></author><summary type="html"><![CDATA[On one coding benchmark, switching from a minimal harness to each model's vendor harness changed scores by as much as +14 points for one model and −13 for another. Benchmark numbers are partly system numbers, and changing the model means re-testing the harness.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://lyr-ai.github.io/assets/img/harness-same-model-two-harnesses.png" /><media:content medium="image" url="https://lyr-ai.github.io/assets/img/harness-same-model-two-harnesses.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">The question mattered more than the model</title><link href="https://lyr-ai.github.io/the-question-mattered-more-than-the-model/" rel="alternate" type="text/html" title="The question mattered more than the model" /><published>2026-09-27T18:00:00-07:00</published><updated>2026-09-27T18:00:00-07:00</updated><id>https://lyr-ai.github.io/the-question-mattered-more-than-the-model</id><content type="html" xml:base="https://lyr-ai.github.io/the-question-mattered-more-than-the-model/"><![CDATA[<p>Jev is a new kind of AI model from a startup called TypeSafe. It doesn’t write text. You
give it some input and a few typed questions (pick one of these options, score this,
is this true?) and it returns answers with probabilities.</p>

<p>It launched with big numbers: “193.6x faster”, “can’t hallucinate”. I skipped those and
read the independent tests instead. Two of them, by different people, point the same way:</p>

<p><strong>How you ask mattered more than which model answered.</strong></p>

<ul>
  <li>Asked one broad question (“is this phishing?”), Jev scored 62.6% and a small ordinary
LLM scored 81.3%. Given the same five narrow questions, both scored 93–95%.</li>
  <li>Asked “which face will this die show?”, Jev was 83% confident and right 19% of the
time. Asked “is it this face?”, it said 19%, close to the true 1 in 6.</li>
</ul>

<figure>
  <img src="/assets/img/jev-one-question-vs-five.png" alt="Slope chart. Asked one question, 'is this phishing?', Jev scores 62.6% and Haiku, a small LLM, 81.3%. Given five narrow questions with weights fitted on 1,000 labelled emails, Jev scores 95.0% and Haiku 93.2%, statistically tied (p = 0.063). A dashed line shows a regex at 91.8%. Jev is about 27 times cheaper and 5 times faster here, at comparable quality." loading="lazy" />
  <figcaption><b>Figure 1.</b> Splitting the question moved both models more than switching models did. On this synthetic dataset a regex already reaches 91.8%, and the five questions were written after reading how the dataset was built.</figcaption>
</figure>

<h2 id="what-jev-is">What Jev is</h2>

<p>Jev’s documentation describes three kinds of question: a <strong>choice</strong> from a list, a
<strong>score</strong> on a rubric, and a true/false question, which returns a probability. Several
questions can go in one request, and each is answered “in parallel and in isolation”.
Input costs $0.042 per million tokens, and output is free.</p>

<p>The docs also say how it’s meant to be used: “Atomic questions, composed in code.” Keep
each question narrow, and combine the answers in your own code.</p>

<p>That advice turns out to be the whole story.</p>

<h2 id="where-would-i-use-it">Where would I use it?</h2>

<p>Vercel, which serves Jev through its AI Gateway, draws the line in one sentence: “Choose
Jev for bounded decisions, code for fixed rules, and a generative model for prose.” In a
real agent or app, that looks like this:</p>

<table>
  <thead>
    <tr>
      <th>Decision point</th>
      <th>Jev?</th>
      <th>Why</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Route a request or ticket to the right team or agent</td>
      <td>Yes</td>
      <td>A choice from a known list</td>
    </tr>
    <tr>
      <td>Review a proposed tool call: run it, or pause for approval</td>
      <td>Yes, as one input</td>
      <td>A yes/no with a probability; the permission itself stays in code</td>
    </tr>
    <tr>
      <td>Score a generated answer against stated requirements</td>
      <td>Yes</td>
      <td>A bounded score</td>
    </tr>
    <tr>
      <td>Write the reply to the user</td>
      <td>No</td>
      <td>“Jev does not generate prose”; use an LLM</td>
    </tr>
    <tr>
      <td>Check a fixed rule, like <code class="language-plaintext highlighter-rouge">balance &gt; 100</code></td>
      <td>No</td>
      <td>Ordinary code is exact and free</td>
    </tr>
    <tr>
      <td>Decide an open-ended plan or strategy</td>
      <td>Probably not</td>
      <td>My judgment: it needs reasoning that doesn’t reduce to a few fixed options</td>
    </tr>
  </tbody>
</table>

<p>The first three rows are uses Vercel lists (alongside prioritizing tickets, categorizing
documents, moderation and choosing which model answers). Vercel also says the guardrail
doesn’t replace permissions: “Enforce access rules and required approvals in application
code before executing a tool.”</p>

<p>The short version: <strong>use it at bounded decision points, not everywhere you’d use an LLM.</strong>
The tests below show how well it holds up there.</p>

<h2 id="test-1-one-question-or-five">Test 1: one question or five</h2>

<p>An independent benchmark (jev-phishing-bench) ran Jev and Claude Haiku 4.5 on 2,000
synthetic emails, half of them phishing.</p>

<p>Asked the single question “is this phishing?”, Jev got <strong>62.6%</strong> right and Haiku
<strong>81.3%</strong>. On that framing, the ordinary LLM clearly wins.</p>

<p>Then the author asked both models five narrow questions (for example, whether the sender
looks generic), and fitted a small logistic regression on 1,000 labelled emails to
combine the answers. Jev reached <strong>95.0%</strong> and Haiku <strong>93.2%</strong>. The benchmark calls that
“statistically tied” (p = 0.063), and Haiku’s AUROC was slightly higher.</p>

<p>The biggest change wasn’t which model answered. It was how the task was decomposed.</p>

<p>What Jev kept was cost and speed. The benchmark’s own summary: “about 27 times cheaper
and 5 times faster than Haiku for signals of comparable quality”. That’s the real case
for a model like this: it makes asking many narrow questions cheap.</p>

<p>Two caveats the benchmark states itself, and which matter:</p>
<ul>
  <li>The emails are synthetic, and a plain regex rule already gets <strong>91.8%</strong>.</li>
  <li>The five questions “were written after reading its URL-evasion taxonomy”. They were
designed knowing how this dataset was built.</li>
</ul>

<h2 id="test-2-which-one-or-is-it-this-one">Test 2: “which one?” or “is it this one?”</h2>

<p>A second independent study (jev-does-not-play-dice) asked Jev about events nobody can
predict, like a fair die roll or a coin flip.</p>

<p>Asked as a <strong>choice</strong>, “which face will it show?”, Jev was badly overconfident. Over 400
die rolls its average confidence was <strong>82.9%</strong> and it was right <strong>19.0%</strong> of the time,
about what guessing gives. On coin flips it said <strong>92%</strong> and was right <strong>52%</strong>.</p>

<p>Asked as a <strong>yes/no</strong>, “is it this face?”, its answers were much closer to the truth: an
average of <strong>19.2%</strong> for an event with a true probability of 16.7%. It isn’t perfect.
For rarer events it drifts high, saying about 15% for something that happens 5% of the
time.</p>

<figure>
  <img src="/assets/img/jev-which-vs-is-it.png" alt="Dot plot. Asked which face a die will show, the model said 82.9% and was right 19% of the time; asked heads or tails, it said 92% and was right 52% of the time. Asked whether a die will show a given face, it said 19.2% against a true 16.7%; for a 1-in-20 event it said 15% against a true 5%." loading="lazy" />
  <figcaption><b>Figure 2.</b> Same uncertainty, asked two ways. The confidence you get back depends on the form of the question.</figcaption>
</figure>

<p>So the same model, facing the same uncertainty, gave very different confidence depending
on how the question was framed.</p>

<p>On a real task the picture can be much better: a separate test on 60 hand-labelled
agent tool calls got 91.7% right, and every wrong answer came with confidence below 1.
Confidence depends on the task and on the question. It has to be checked, not trusted.</p>

<h2 id="cant-hallucinate-means-cant-leave-the-schema">“Can’t hallucinate” means “can’t leave the schema”</h2>

<p>TypeSafe says Jev “can’t hallucinate”. Its own launch post qualifies this: “Our number is
not empirical. Schema matching is guaranteed.”</p>

<p>That’s the precise meaning. Jev can only answer with the options you gave it. It can’t
invent a category or return malformed output. But it can still pick the wrong option,
confidently, as the die shows.</p>

<h2 id="what-id-take-from-this">What I’d take from this</h2>

<p>If you use a decision model like Jev:</p>

<ul>
  <li><strong>Design questions, not prompts.</strong> Narrow questions combined in code did better than
one broad one, for both models tested.</li>
  <li><strong>Its advantage is cost and speed, not accuracy.</strong> On this evidence, it makes the
decomposition affordable rather than making each answer smarter.</li>
  <li><strong>Check its confidence on your own labelled data</strong> before acting on it. The form of the
question changed it.</li>
</ul>

<h2 id="what-i-dont-know">What I don’t know</h2>

<ul>
  <li>Both studies are small and single-author; one uses synthetic data that a regex handles
well.</li>
  <li>Nothing here measures production traffic.</li>
  <li>TypeSafe’s speed and cost multipliers (“193.6x faster, 444.6x cheaper”) are its own,
from its own workflows, and it says they are “on the higher end of real world gains”.
I haven’t used them.</li>
</ul>

<hr />

<p><em>Sources: TypeSafe’s launch post and documentation; Vercel’s guides “When should you use
Jev instead of a chat model?”, “7 practical Jev use cases” and “Where does Jev fit in an
AI agent loop?”; jev-phishing-bench
(github.com/anisselbd/jev-phishing-bench); jev-does-not-play-dice
(github.com/KantaHayashiAI/jev-does-not-play-dice); the tool-call risk benchmark by
webofmike on dev.to. Each number was read on the original page or repository.</em></p>]]></content><author><name>Ruxi Zhang</name></author><summary type="html"><![CDATA[Jev, a new "decision model", launched with big claims. Two independent tests point somewhere more useful: how you ask changed the result more than which model answered.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://lyr-ai.github.io/assets/img/jev-one-question-vs-five.png" /><media:content medium="image" url="https://lyr-ai.github.io/assets/img/jev-one-question-vs-five.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">The engine kept moving underneath Ollama</title><link href="https://lyr-ai.github.io/the-engine-kept-moving-underneath-ollama/" rel="alternate" type="text/html" title="The engine kept moving underneath Ollama" /><published>2026-09-27T17:30:00-07:00</published><updated>2026-09-27T17:30:00-07:00</updated><id>https://lyr-ai.github.io/the-engine-kept-moving-underneath-ollama</id><content type="html" xml:base="https://lyr-ai.github.io/the-engine-kept-moving-underneath-ollama/"><![CDATA[<p>Should a project build its own inference engine, or rely on an upstream one?</p>

<p>Ollama’s git history gives an unexpected answer: both.</p>

<p>Over three years, Ollama moved from borrowing llama.cpp, to copying it into its own
repository, to building its own GGML-based engine. Then, in May 2026, it deleted that
engine and handed GGUF models to upstream llama.cpp’s server.</p>

<p>That sounds like a reversal. It wasn’t quite.</p>

<p>At the same time, Ollama was building another engine of its own, on Apple’s MLX, and it
kept that one. By September 2026 the stack had split:</p>

<ul>
  <li><strong>GGUF models → upstream llama.cpp.</strong> Ollama borrows it.</li>
  <li><strong>safetensors models → Ollama’s MLX engine.</strong> Ollama builds it.</li>
</ul>

<p>The interesting decision wasn’t <em>build or borrow</em>. It was <strong>where to own the engine</strong>.</p>

<p>What makes this easy to miss is that users never had to notice: every command and API
endpoint a 2023 user had is still there. I read Ollama’s full git history (5,796 commits,
June 2023 to late September 2026) to see what happened underneath.</p>

<figure>
  <img src="/assets/img/systems-seen-ollama-engine.png" alt="A timeline from 2023 to 2026. On top, a green band labelled 'what users run' that only gets wider, at chat, OpenAI-compatible /v1, embed and accounts. Below, the inference engine moves between four levels of ownership: borrowed in a separate process, borrowed in-process, copied into the repo, and built by Ollama. It ends in 2026 split in two: upstream llama-server for GGUF models, and Ollama's own MLX engine for safetensors." loading="lazy" />
  <figcaption><b>Figure 1.</b> Where Ollama's inference engine lived, 2023–2026. <a href="/systems-seen/ollama/">Open the interactive version</a> to step through it.</figcaption>
</figure>

<h2 id="what-users-saw">What users saw</h2>

<p>The surface only grew. <code class="language-plaintext highlighter-rouge">ollama run MODEL</code> and the eight original endpoints
(<code class="language-plaintext highlighter-rouge">/api/generate</code>, <code class="language-plaintext highlighter-rouge">/api/pull</code>, <code class="language-plaintext highlighter-rouge">/api/tags</code> and the rest) exist at every point I
checked, including today. On top of them came <code class="language-plaintext highlighter-rouge">/api/chat</code> (December 2023), an
OpenAI-compatible API (February 2024), <code class="language-plaintext highlighter-rouge">/api/embed</code> (July 2024) and account endpoints
(September 2025).</p>

<p>Nothing a 2023 user relied on was removed. From the perspective of that original
surface, Ollama remained compatible while adding more.</p>

<h2 id="what-moved-underneath">What moved underneath</h2>

<p>The engine is a different story. Reading the tree and the commit messages, it went
through these places:</p>

<ul>
  <li><strong>July 2023: Go bindings.</strong> Ollama called llama.cpp in-process (“add llama.cpp go
bindings”).</li>
  <li><strong>August 2023: llama.cpp’s server, as a separate process.</strong> “subprocess llama.cpp
server” removed the C code; Ollama now ran llama.cpp’s own server and talked to it.
From January 2024 it carried its own patches on top.</li>
  <li><strong>October 2024 (v0.4.0): llama.cpp copied into the repo.</strong> “Remove submodule and
shift to Go server”: llama.cpp was vendored, and driven by a Go server.</li>
  <li><strong>December 2024 onwards: Ollama’s own engine.</strong> “Runner for Ollama engine”, then a
GGML-based backend and model implementations written in Go, running alongside the
vendored llama.cpp.</li>
</ul>

<p>Plotted by how much of the engine Ollama owned, that’s a steady climb: from borrowing
llama.cpp, to copying it in, to building its own.</p>

<h2 id="it-looks-like-build-then-borrow">It looks like build, then borrow</h2>

<p>Then, on 29 May 2026, one commit reversed most of it:</p>

<blockquote>
  <p>Remove the vendored GGML and llama.cpp backend, CGO runner, Go model implementations,
and sample. llama-server (built from upstream llama.cpp via FetchContent) is now the
sole inference engine for GGUF-based models.</p>
</blockquote>

<p>The engine Ollama had built for GGUF models was deleted. So was the copied llama.cpp.
Ollama went back to running llama.cpp’s server as a separate program, this time
upstream’s own, built at build time, with a small compatibility layer so it can load
Ollama-format model files.</p>

<p>It is tempting to read that as a verdict: they tried to build their own engine, and it
didn’t work out.</p>

<h2 id="thats-the-wrong-conclusion">That’s the wrong conclusion</h2>

<p>The same commit carries a parenthesis:</p>

<blockquote>
  <p>(Safetensor based models continue to run on the new MLX engine.)</p>
</blockquote>

<p>Ollama had been building a second engine of its own, on Apple’s MLX, since February</p>
<ol>
  <li>About three and a half months after the GGUF engine was removed, the MLX one was
promoted: “the only
Go inference runner left and is no longer experimental.”</li>
</ol>

<p>So in 2026 Ollama didn’t stop building. It split:</p>

<table>
  <thead>
    <tr>
      <th>Model format</th>
      <th>Engine</th>
      <th>Ollama’s relationship to it</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>GGUF</td>
      <td>upstream llama.cpp server</td>
      <td>borrows it</td>
    </tr>
    <tr>
      <td>safetensors</td>
      <td>Ollama’s MLX engine</td>
      <td>builds it</td>
    </tr>
  </tbody>
</table>

<h2 id="the-decision-wasnt-build-or-borrow">The decision wasn’t build or borrow</h2>

<p>“Build or buy” is usually framed as one decision for a whole component. Ollama’s history
shows it made twice, and the second time it wasn’t made for the whole component at all.
The line ended up being drawn by model format: GGUF moved to upstream llama.cpp, while
safetensors stayed on Ollama’s own MLX engine.</p>

<p>The more useful question isn’t “should we build our own engine?” It’s “<strong>where should
we own the engine?</strong>”</p>

<p>The only reason the maintainers give is in that commit: “This allows us to more rapidly
pick up new capabilities and fixes from llama.cpp as they come out.”</p>

<h2 id="what-this-doesnt-tell-us">What this doesn’t tell us</h2>

<p>Everything above comes from the repository itself: the file tree at different dates,
and the commit messages. So there is a lot it can’t say:</p>

<ul>
  <li>whether the switches made Ollama faster or slower, or changed anything users noticed;</li>
  <li>what the in-house GGML engine cost to maintain, and why exactly it was dropped,
beyond the one sentence above;</li>
  <li>how the team made these decisions;</li>
  <li>whether the MLX engine will one day go the same way.</li>
</ul>

<p>Commit messages are the maintainers’ own framing. I’ve quoted them, and checked each
phase against the tree, but I haven’t filled the gaps with guesses.</p>

<p>What the history does show is a pattern worth recognising in any system you depend on:
a surface that never broke, over an implementation boundary that kept moving.</p>

<hr />

<p><em>Method: a full clone of github.com/ollama/ollama (HEAD <code class="language-plaintext highlighter-rouge">16b4376a</code>, 2026-09-26), read
from the raw history; no third-party summaries. Engine phases: <code class="language-plaintext highlighter-rouge">6093a88c</code>, <code class="language-plaintext highlighter-rouge">42998d79</code>,
<code class="language-plaintext highlighter-rouge">b754f5a6</code>, <code class="language-plaintext highlighter-rouge">ed443a03</code>, <code class="language-plaintext highlighter-rouge">dcfb7a10</code>, <code class="language-plaintext highlighter-rouge">d8cc798c</code>, <code class="language-plaintext highlighter-rouge">9db4bdba</code>, <code class="language-plaintext highlighter-rouge">2e036e7c</code>. Surface
additions: <code class="language-plaintext highlighter-rouge">7a0899d6</code> (<code class="language-plaintext highlighter-rouge">/api/chat</code>), <code class="language-plaintext highlighter-rouge">453f572f</code> (<code class="language-plaintext highlighter-rouge">/v1</code>), <code class="language-plaintext highlighter-rouge">b9f5e16c</code> (<code class="language-plaintext highlighter-rouge">/api/embed</code>),
<code class="language-plaintext highlighter-rouge">8b894933</code> (accounts).</em></p>]]></content><author><name>Ruxi Zhang</name></author><summary type="html"><![CDATA[Ollama didn't choose between building and borrowing its inference engine. By 2026 it borrowed upstream llama.cpp for GGUF models and built its own MLX engine for safetensors.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://lyr-ai.github.io/assets/img/systems-seen-ollama-engine.png" /><media:content medium="image" url="https://lyr-ai.github.io/assets/img/systems-seen-ollama-engine.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">What happened when I made agent memory transitions explicit</title><link href="https://lyr-ai.github.io/what-happened-when-i-made-memory-transitions-explicit/" rel="alternate" type="text/html" title="What happened when I made agent memory transitions explicit" /><published>2026-09-27T12:00:00-07:00</published><updated>2026-09-27T12:00:00-07:00</updated><id>https://lyr-ai.github.io/what-happened-when-i-made-memory-transitions-explicit</id><content type="html" xml:base="https://lyr-ai.github.io/what-happened-when-i-made-memory-transitions-explicit/"><![CDATA[<p>Two weeks ago, in <a href="/from-retrieval-to-state/">From retrieval to state</a>,
I argued that long-lived agent memory should be modeled as state and the
transitions between states, not only as retrieval.</p>

<p>Then I implemented those transitions in TypedMem and ran a complete scenario
through them: an agent investigating a production incident. Two things became
clearer than they were in that post.</p>

<p><strong>First, “new information arrived” is not one memory operation.</strong></p>

<p><strong>Second, recording how memory changed is not the same as recording why an agent
changed its mind.</strong> I had blurred that distinction myself.</p>

<h2 id="one-write-several-possible-meanings">One write, several possible meanings</h2>

<p>An operations agent is investigating why clients can’t connect to a logging
service. Information arrives from runtime telemetry, support cases, a
certificate validation system, and later telemetry.</p>

<p>The scenario uses three kinds of memory, each with its own transition rule:</p>

<table>
  <thead>
    <tr>
      <th>Memory type</th>
      <th>When another record hits the same subject</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">evidence</code></td>
      <td><strong>reinforce</strong>: one record, more sources</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">assessment</code></td>
      <td><strong>flag</strong>: keep both, mark them as conflicting</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">decision</code></td>
      <td><strong>supersede</strong>: keep the old one in history, the new one is current</td>
    </tr>
  </tbody>
</table>

<p>Here is what happened in the run:</p>

<ul>
  <li><strong>Telemetry</strong> reports that certificate errors are associated with the
failures, and the agent stores it as <code class="language-plaintext highlighter-rouge">evidence</code>. <strong>Support cases</strong> report the
same thing. TypedMem doesn’t create a second record: it keeps <strong>one record
with two sources</strong>, and its confidence rises from 0.60 to 0.72.</li>
  <li><strong>The agent</strong> stores its own <code class="language-plaintext highlighter-rouge">assessment</code>: the certificate is causing the
failures. <strong>The validation system</strong> reports that the certificate is valid,
stored as an <code class="language-plaintext highlighter-rouge">assessment</code> on the same subject. Both are kept and marked as
conflicting. The conflict <strong>stays open</strong>: nothing later resolves it, and the
memory layer doesn’t pretend otherwise.</li>
  <li><strong>The agent decides</strong> to investigate the certificate. Later, telemetry shows
that failures correlate with lost connectivity to the logging service, and the
agent revises its decision. The new decision <strong>supersedes</strong> the old one: the
old one stays in history, and only the new one is current.</li>
</ul>

<figure>
  <img src="/assets/img/memory-transitions-filing.png" alt="New information arrives and the application files it by type and subject; TypedMem doesn't read the text. Filed as evidence, it reinforces: one record with two sources (telemetry, support), confidence 0.60 to 0.72. Filed as an assessment, it is flagged: the agent's 'causing it' and validation's 'it's valid' are both kept, marked as conflicting, still unresolved. Filed as a decision, it supersedes: the connectivity decision is current and the certificate decision is kept as history. Below, the wrong turn: the same 'certificate is valid' filed as evidence is merged as a second supporting source, one record, confidence 0.60 to 0.72, no conflict." loading="lazy" />
  <figcaption><b>Figure 1.</b> The same input can become a conflict or support, depending on how the application files it.</figcaption>
</figure>

<h3 id="typedmem-didnt-discover-any-of-these-relationships">TypedMem didn’t discover any of these relationships</h3>

<p>The text never decided anything. I did, when I told the system which kind of
memory each piece of information was:</p>
<ul>
  <li>telemetry and support cases are <code class="language-plaintext highlighter-rouge">evidence</code>;</li>
  <li>the two judgments about the certificate are <code class="language-plaintext highlighter-rouge">assessment</code>s about the same
subject;</li>
  <li>the investigation path is a <code class="language-plaintext highlighter-rouge">decision</code>.</li>
</ul>

<p>TypedMem then applied the rule each type declares. It doesn’t read the text: it
matches records by type and subject, and the type’s rule decides what happens.</p>

<p>While writing this I ran the obvious mistake. I filed the validation result as
another <code class="language-plaintext highlighter-rouge">evidence</code> record on the same subject as the agent’s assessment. There was
no conflict. The store kept <strong>one</strong> record, whose text was still “The certificate
is causing client failures”. The validation system became its <strong>second
supporting source</strong>, and confidence rose from 0.60 to 0.72. The disagreement
disappeared into agreement.</p>

<p>That isn’t a bug in the policy engine. It shows the boundary: <strong>the memory layer
can enforce semantics, but the application has to assign the right semantics
first.</strong></p>

<h2 id="authority-is-a-policy-not-the-truth">Authority is a policy, not the truth</h2>

<p>In an early outline of this post I wrote that picking the more trusted source
“silently turns authority into a truth oracle”. That was too strong.</p>

<p>Sometimes authority is exactly the rule you want. For TypedMem’s replace-style
memory types, a newer record from a weaker source can’t displace an existing one
from a stronger source: it is ignored. In an <a href="https://github.com/lyr-ai/reliagent-bench/blob/measure/states-0.9.3/external/results/states-0.9.3.md">earlier measurement</a>,
typed memory with that rule scored 1.00 on the authority cases.</p>

<p>Other kinds of memory want disagreement to stay visible. That’s what the flag rule
is for. And evidence wants agreement to accumulate.</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>        the same incoming information
                     │
                     ▼
       depends on the memory's contract
         ┌───────────┼───────────┐
         ▼           ▼           ▼
       veto         flag      reinforce
</code></pre></div></div>

<p>The point isn’t that authority is bad. It’s that different kinds of memory need
different transition rules, not one universal resolver of truth.</p>

<h2 id="what-the-event-log-recorded">What the event log recorded</h2>

<p>Every transition is recorded. Grouped by what it means, the run’s history is:</p>

<ol>
  <li>evidence added (telemetry)</li>
  <li>evidence reinforced (support cases)</li>
  <li>assessment added (the agent: the certificate is causing failures)</li>
  <li>competing assessment stored; both records flagged (validation)</li>
  <li>decision added (investigate the certificate)</li>
  <li>evidence added (connectivity)</li>
  <li>decision superseded (investigate connectivity)</li>
</ol>

<p>The raw log is more detailed. Flagging and superseding touch both records, so they
write paired events: nine events for these seven steps.</p>

<p>Replaying the whole log rebuilds the store exactly: the same six records. Replaying
only the events before step 7 rebuilds the earlier state, in which the current
decision was still “investigate the certificate”.</p>

<h2 id="memory-history-is-not-reasoning-history">Memory history is not reasoning history</h2>

<p>This is the part I got wrong.</p>

<p>The main figure of the first post had this caption:</p>

<blockquote>
  <p>The important object is not an isolated memory item. It is the transition
structure that explains how the present emerged from the past.</p>
</blockquote>

<p>I would phrase that differently now. Here is what the implementation actually
knows:</p>

<table>
  <thead>
    <tr>
      <th>Kind of provenance</th>
      <th>Question</th>
      <th>Recorded?</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>Source</strong></td>
      <td>Where did this record come from?</td>
      <td>Yes</td>
    </tr>
    <tr>
      <td><strong>Transition</strong></td>
      <td>What happened to stored memory?</td>
      <td>Yes</td>
    </tr>
    <tr>
      <td><strong>Reasoning</strong></td>
      <td>Which evidence caused this decision?</td>
      <td>Only if the application records it</td>
    </tr>
  </tbody>
</table>

<p>The event log can tell me that the connectivity evidence was stored before the
investigation decision changed. It cannot tell me that the evidence <strong>caused</strong>
the change. In the run, the new decision’s only source is the agent, and nothing
links it to the evidence.</p>

<p><strong>Time order is not causal explanation.</strong></p>

<p>It’s tempting to draw the investigation as one chain (telemetry → hypothesis →
conflict → new evidence → decision revised). That drawing would claim knowledge
the system never stored.</p>

<p>Replay has the same boundary. It answers “what did memory contain at this point?”,
not “why did the agent decide that?”</p>

<figure>
  <img src="/assets/img/memory-provenance-kinds.png" alt="Three kinds of provenance. Source (where did this record come from?) and transition (what happened to stored memory?) are recorded, as sources on every record and as the event log and replay. Reasoning (which evidence led to this decision?) is not recorded unless the application writes it. Below, steps 6 and 7 of the run: evidence added (failures correlate with lost connectivity), then the decision superseded (investigate connectivity), with no link stored between them. Time order is not causal explanation." loading="lazy" />
  <figcaption><b>Figure 2.</b> Two of the three are recorded. Steps 6 and 7 are adjacent in the log, but nothing links them: the new decision's only source is the agent.</figcaption>
</figure>

<h2 id="where-the-memory-layer-ends">Where the memory layer ends</h2>

<p>I started with the idea that long-lived memory needs explicit state
transitions. Implementing it made the boundary sharper:</p>

<ul>
  <li><strong>The application assigns meaning.</strong></li>
  <li><strong>The memory layer enforces the declared transition.</strong></li>
  <li><strong>The event log records what changed.</strong></li>
  <li><strong>None of these, by itself, records why the agent reasoned from A to B.</strong></li>
</ul>

<p>That leaves the question I don’t know the answer to yet. How much decision
provenance belongs in a general memory layer? Should links from evidence to
decisions be first-class memory structure, or stay application-specific?</p>

<p>If you’ve built long-running agents and had to answer “why did it decide
that?”, I’d like to know where you ended up putting that link.</p>

<hr />

<p><em>The scenario is a runnable example,
<a href="https://github.com/lyr-ai/typedmem/blob/main/examples/incident_investigation.py"><code class="language-plaintext highlighter-rouge">examples/incident_investigation.py</code></a>.
It uses TypedMem’s public API, and a test pins its output, so the run described
here is the run you get.</em></p>]]></content><author><name>Ruxi Zhang</name></author><summary type="html"><![CDATA[New information isn't one memory operation, and memory history isn't reasoning history. What running a complete scenario through explicit memory transitions taught me about where the memory layer ends.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://lyr-ai.github.io/assets/img/memory-transitions-filing.png" /><media:content medium="image" url="https://lyr-ai.github.io/assets/img/memory-transitions-filing.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Why agent CI can’t be treated like deterministic tests</title><link href="https://lyr-ai.github.io/why-agent-ci-cant-be-treated-like-deterministic-tests/" rel="alternate" type="text/html" title="Why agent CI can’t be treated like deterministic tests" /><published>2026-09-26T17:56:13-07:00</published><updated>2026-09-26T17:56:13-07:00</updated><id>https://lyr-ai.github.io/why-agent-ci-cant-be-treated-like-deterministic-tests</id><content type="html" xml:base="https://lyr-ai.github.io/why-agent-ci-cant-be-treated-like-deterministic-tests/"><![CDATA[<p>Continuous integration rests on one assumption so basic we rarely say it out
loud: <strong>run the same code on the same input and you get the same result.</strong> A
test that passed on <code class="language-plaintext highlighter-rouge">main</code> and fails on your branch is evidence against your
branch. That assumption is what makes a red check mean something.</p>

<p>AI agents break it.</p>

<p>Run an agent twice on the same task, with the same code, model and prompt, and
it can read different files, call tools in a different order and end
somewhere else. Sometimes it succeeds both times, sometimes once. So when a
pull request touches the agent and the eval score moves, what should the
check believe?</p>

<h2 id="a-check-that-failed-for-the-right-reason">A check that failed for the right reason</h2>

<p>Here is a pull request that looks like a harmless refactor. It changes one
line in a checkout handler, and the field it now reads, <code class="language-plaintext highlighter-rouge">amount_cents</code>, does not
exist.</p>

<figure>
  <img src="/assets/img/agentseism-example-pr-regression.png" alt="A GitHub pull request, 'checkout: use amount_cents for the receipt amount'. The github-actions bot has posted an AgentSeism report headed 'REGRESSION — do not merge without review'. Its scorecard has two rows. Task success (broad): baseline 0.93, PR 0.68, change -0.25 with interval -0.52 to -0.04, decision PASS. Task success (capability): 7 of 7 tasks monitored, 1 collapsed, checkout 8/8 to 0/8, decision REGRESSION. The evidence list shows checkout fired with p=0.0001 and a non-blocking warning on return-label, 7/8 to 3/8." loading="lazy" />
  <figcaption><b>Figure 1.</b> <a href="https://github.com/lyr-ai/agentseism-example/pull/2">A public pull request</a> in an example repository. The agent there is simulated: real handler code, fixed per-task success rates, and outcomes that vary from run to run. The CI check around it is the real one.</figcaption>
</figure>

<p>The check ran each of seven tasks eight times on <code class="language-plaintext highlighter-rouge">main</code> and eight times on the
branch. Overall success fell from <strong>0.93 to 0.68</strong>, and two things happened that
ordinary CI has no vocabulary for:</p>

<ul>
  <li><strong>The broad reliability check passed.</strong> The evidence supported some
deterioration, but not a suite-wide drop of at least the 10-point threshold
declared in advance.</li>
  <li><strong>The check still failed</strong>, because <code class="language-plaintext highlighter-rouge">checkout</code>, which had succeeded <strong>8 of 8</strong>
times on <code class="language-plaintext highlighter-rouge">main</code>, succeeded <strong>0 of 8</strong> times on the branch.</li>
</ul>

<p>A third detail is easy to miss: <code class="language-plaintext highlighter-rouge">return-label</code> went from 7/8 to 3/8 and was
flagged as a <em>warning</em>, not a failure. Its success rate didn’t change in this
PR, so that drop is pure run-to-run wobble, and a warning is the right amount
of alarm for it.</p>

<p>A single pass/fail test cannot express any of that. Neither can a single eval
score.</p>

<h2 id="one-run-is-a-sample-not-a-replay">One run is a sample, not a replay</h2>

<p>A deterministic test is a <em>replay</em>: run it again and you learn nothing new. An
agent run is a <em>sample</em> from a distribution of possible runs. That changes what
a comparison between two branches can tell you, in three ways.</p>

<figure>
  <img src="/assets/img/agentseism-three-cases.svg" alt="Three cards. Normal noise: task success 91% to 86%, PASS, ordinary run-to-run wobble. Capability collapse: 91% to 77%, REGRESSION, checkout 8/8 to 0/8. Not enough evidence: 100% to 83%, NEED EVIDENCE, 3 tasks with 2 runs per side." loading="lazy" />
  <figcaption><b>Figure 2.</b> The three situations a stochastic check has to tell apart. These come from <code>seism demo</code>, which feeds fixed, stated outcomes through the real decision engine. They illustrate the cases and are not evidence.</figcaption>
</figure>

<h3 id="1-scores-move-when-nothing-changed">1. Scores move when nothing changed</h3>

<p>On a real coding agent (mini-swe-agent with Claude Haiku 4.5, on five SWE-bench
tasks run five times each), we ran an <strong>unchanged</strong> candidate against its own
baseline. Success went from <strong>0.92 to 0.88</strong>. One task went from 3/5 to 2/5 with
no code change at all.</p>

<p>A deterministic mindset reads that as “the PR made it worse”. It didn’t. There
was no PR. A check that blocks on this trains developers to click “re-run”
until it goes green, and after that nobody trusts it.</p>

<h3 id="2-averages-hide-breakage">2. Averages hide breakage</h3>

<p>The opposite failure is quieter. In the pull request above, the average fell 25
points, but not evenly: most of it was one task collapsing completely. Averaged
over the suite, a total failure of one capability looks like a moderate,
uncertain drop.</p>

<figure>
  <a href="/agentseism/explorer/?s=collapse"><img src="/assets/img/agentseism-explorer-preview.png" alt="A measurement plate of seven capabilities, each with eight runs on main and eight on the pull request. Standing marks are successful runs; dots on the floor are failures. Six capabilities barely change. Checkout goes from eight standing marks on main to eight red dots on the pull request, 8/8 to 0/8. Beneath, a seismic trace shows small tremors across the suite and then a sharp fall at checkout. The average fell 14 points; checkout fell 100." loading="lazy" /></a>
  <figcaption><b>Figure 3.</b> The average fell 14 points. One capability disappeared.
  <a href="/agentseism/explorer/?s=collapse"><b>▶ Explore what happened</b></a>: an interactive version of the <code>seism demo</code> collapse scenario, where every mark is one run.</figcaption>
</figure>

<p>This isn’t hypothetical for us. The first version of our own check measured only
the suite-wide average, and on a real agent it passed a change that cut success
from 0.92 to 0.52, with two tasks collapsing. The statistics were computed
correctly. They were answering the wrong question. That failure deserves its own
post.</p>

<h3 id="3-sometimes-the-honest-answer-is-not-enough-evidence">3. Sometimes the honest answer is “not enough evidence”</h3>

<p>Three tasks, run twice each, going from 100% to 83% is not a regression and not
a pass. It is too little data. A binary pass/fail setup has no honest way to
represent that, so it gets rounded to whichever side of the threshold it falls
on.</p>

<h2 id="why-not-just-run-your-eval-more-times">Why not just run your eval more times?</h2>

<p>Running more repetitions is the obvious fix, and it helps, but it doesn’t
answer the question on its own. More numbers still leave four decisions open:</p>

<ul>
  <li><strong>What counts as independent evidence?</strong> Fifty reruns of one task tell you a
lot about that task and almost nothing about the others. For a claim about
the whole suite, the task is the unit, not the run.</li>
  <li><strong>How big a change matters?</strong> A 2-point drop can be real and still not worth
blocking a merge over. You have to say how much you care about before
looking.</li>
  <li><strong>How do you avoid false alarms across many capabilities?</strong> Watch 50 tasks
with a simple per-task rule (“block if a task that passed at least 7/8 drops
to at most 2/8”) and some healthy PR will eventually trip it by chance. In our
simulations of a flaky agent, that rule false-blocks about 1.4% of unchanged
PRs at 5 tasks and <strong>13.4% at 50</strong>. A rule that corrects for the number of
tasks stays under 1%
(<a href="https://github.com/lyr-ai/agentseism/blob/master/analysis/CI_V1_BASELINE_CHALLENGE.md">computation</a>,
table C6).</li>
  <li><strong>What do you do when there isn’t enough data?</strong> You need a third answer
besides pass and fail.</li>
</ul>

<p>Repeating the eval gives you the evidence. Something still has to turn that
evidence into a merge decision.</p>

<h2 id="turning-repeated-runs-into-a-merge-decision">Turning repeated runs into a merge decision</h2>

<p>That is what <a href="https://github.com/lyr-ai/agentseism">AgentSeism</a> does. You give
it a command that runs your agent on a task and one that checks the result. It
runs your base branch and your PR several times each, keeps each task’s results
together, and returns one of four verdicts:</p>

<table>
  <thead>
    <tr>
      <th>Verdict</th>
      <th>Meaning</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><strong>PASS</strong></td>
      <td>Neither gate found enough evidence of a material regression.</td>
    </tr>
    <tr>
      <td><strong>REGRESSION</strong></td>
      <td>The suite as a whole is confidently worse by at least your threshold, <strong>or</strong> a task that reliably worked has collapsed.</td>
    </tr>
    <tr>
      <td><strong>INSUFFICIENT EVIDENCE</strong></td>
      <td>Too little data to decide. It blocks the merge, but it is not reported as a regression.</td>
    </tr>
    <tr>
      <td><strong>INCOMPARABLE</strong></td>
      <td>The model, runtime or dependencies differ between the two sides, so nothing is compared.</td>
    </tr>
  </tbody>
</table>

<p>The two gates behind REGRESSION answer different questions, which is why they
are separate: <em>did the suite get broadly worse?</em> and <em>did something that used to
work stop working?</em> The pull request above is the case where the answers
differ.</p>

<p>One correction to <a href="/not-every-behavioral-change-is-a-regression/">the previous note in this series</a>.
It ended on “Comparability first. Outcomes decide. Traces explain.” The first
two became the product. The third, localising a regression to a stage of the
agent’s trajectory, did not, and AgentSeism does not claim it.</p>

<h2 id="how-much-should-you-trust-it">How much should you trust it?</h2>

<p>We froze the decision rule before testing it, then ran it on seven SWE-bench
tasks it had never seen, with a pre-registered protocol: 224 real agent runs, 0
invalid, $38.49 in API cost. It passed an unchanged agent, and it caught both
degradations that had been predicted in advance, a capability collapse and a
broad collapse.</p>

<p>That evidence is narrow, and it is worth saying how:</p>

<ul>
  <li><strong>One agent, one model, one kind of regression.</strong> All of it is mini-swe-agent
with Claude Haiku 4.5, and every degradation was made by cutting the agent’s
step budget.</li>
  <li><strong>Moderate regressions are hard to catch.</strong> The rule is built to avoid
blocking healthy PRs, so a task falling from 100% to 75% usually won’t block
a merge.</li>
  <li><strong>It isn’t cheap yet.</strong> Eight runs per task on each side is a validation
design, not an optimised CI budget.</li>
  <li><strong>No one outside the project has used it yet.</strong></li>
</ul>

<p>The <a href="https://github.com/lyr-ai/agentseism/blob/master/analysis/ci_v1/stageC/RESULTS.md">full protocol and results</a>,
including what the study did <em>not</em> establish, are in the repository.</p>

<h2 id="try-it">Try it</h2>

<p>If your agent’s eval score has ever moved on a PR and you couldn’t tell whether
to believe it:</p>

<div class="language-bash highlighter-rouge"><div class="highlight"><pre class="highlight"><code>git clone https://github.com/lyr-ai/agentseism.git <span class="o">&amp;&amp;</span> <span class="nb">cd </span>agentseism
python3.11 <span class="nt">-m</span> venv .venv <span class="o">&amp;&amp;</span> <span class="nb">source</span> .venv/bin/activate
pip <span class="nb">install</span> <span class="nt">-e</span> <span class="nb">.</span>
seism demo
</code></pre></div></div>

<p>It takes about a second, with no API key and no Docker. Then look at the two
public pull requests in
<a href="https://github.com/lyr-ai/agentseism-example/pulls">agentseism-example</a>: one
docs-only change that passes, and the one-line bug above.</p>

<p>Or run your own eval eight times on an unchanged branch, and see how stable it
really is.</p>

<p>If this is a problem you have too, a ⭐ on
<a href="https://github.com/lyr-ai/agentseism">the repo</a> tells me it’s worth continuing.
An issue telling me where it breaks on your agent is worth even more.</p>]]></content><author><name>Ruxi Zhang</name></author><summary type="html"><![CDATA[Run the same agent twice and you get two different results. So what should a pull request check believe? A 25-point drop that wasn't enough evidence on its own, a capability that failed every run, and why running your eval more times doesn't settle it.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://lyr-ai.github.io/assets/img/agentseism-social-preview.png" /><media:content medium="image" url="https://lyr-ai.github.io/assets/img/agentseism-social-preview.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Not every behavioral change is a regression</title><link href="https://lyr-ai.github.io/not-every-behavioral-change-is-a-regression/" rel="alternate" type="text/html" title="Not every behavioral change is a regression" /><published>2026-09-20T09:00:00-07:00</published><updated>2026-09-20T09:00:00-07:00</updated><id>https://lyr-ai.github.io/not-every-behavioral-change-is-a-regression</id><content type="html" xml:base="https://lyr-ai.github.io/not-every-behavioral-change-is-a-regression/"><![CDATA[<blockquote>
  <p><strong>Update, September 2026.</strong> AgentSeism’s product direction changed after this
note. Comparability and outcome-based CI decisions became the product. Trace
localisation, the “traces explain” step and the root-cause stage in Figure 2,
remains research-era work and is not part of AgentSeism’s product today. See
<a href="/why-agent-ci-cant-be-treated-like-deterministic-tests/">Method note 03</a>.</p>
</blockquote>

<p>Software regression testing assumes repeated executions are reasonably stable.
If the same test passes before a change and fails afterward, the change is a
plausible cause.</p>

<p>AI agents break this assumption.</p>

<p>An agent can take a different path even when its code, task, and model
configuration appear unchanged. It may inspect different files, call tools in a
different order, write a different patch, or recover from an error through a
different route. Some of these changes are harmful. Many are not.</p>

<p>This creates a basic problem for agent development:</p>

<blockquote>
  <p>When an agent behaves differently after a change, how do we determine whether
the change is a real regression, harmless stochastic variation, or the result
of an incomparable execution environment?</p>
</blockquote>

<p>I built <a href="https://github.com/lyr-ai/agentseism">AgentSeism</a> to explore this
question. The project began as a framework for detecting behavioral instability
in agent trajectories. Experiments changed that direction. They showed that
behavioral difference alone is not a reliable release signal.</p>

<p>The more useful design is an outcome-grounded regression layer for stochastic
agents:</p>

<ol>
  <li>Check whether the two evaluations are comparable.</li>
  <li>Decide whether task outcomes regressed.</li>
  <li>Use traces only to diagnose a regression after it has been established.</li>
</ol>

<p>That ordering matters.</p>

<figure>
  <img src="/assets/img/agentseism-two-failure-modes.png" alt="Two panels. Left: five bars showing steps taken by five independent runs of the same coding task at temperature zero — 31, 41, 33, 47 and 31 steps, with final patch sizes 843, 1229, 978, 1939 and 690 bytes. The first four bars are green, marked resolved; the fifth is red, not resolved. A composite trace detector fires on all six pairs of correct runs. Right: a horizontal bar chart of 23 archived fork roots. Structured actions differ on 23 of 23; zero match. The agent was held completely fixed and only the GPU and driver changed, from an A100-SXM4-80GB to an H100 PCIe." loading="lazy" />
  <figcaption><b>Figure 1.</b> Two failure modes, both measured on frozen artifacts. Left: trajectory
  consistency would have rejected correct work. Right: an environment change looks exactly like an agent change.</figcaption>
</figure>

<h2 id="different-trajectories-can-all-be-correct">Different trajectories can all be correct</h2>

<p>Consider several independent runs of the same coding task under the same agent
configuration.</p>

<p>In one frozen batch, five runs produced five different final workspace states.
Four of the five solutions were accepted by the task evaluator; one was not.
The four correct runs took 31, 41, 33 and 47 steps and submitted patches of
843, 1229, 978 and 1939 bytes. No two shared a tool-usage profile.</p>

<p>This small example exposes both sides of the problem.</p>

<p><strong>Trajectory consistency would have been too strict.</strong> The four successful runs
did not converge on one canonical trajectory or one identical final state. A
test that required them to reproduce the same commands, intermediate files, or
patch structure would have rejected valid solutions. I ran the check to see how
badly: applying the kind of distribution-shift thresholds a behavioral
fingerprint uses — trace length, tool mix, patch size — a composite detector
fires on <strong>all six</strong> pairs of correct runs. Ground truth: none of the six is a
regression.</p>

<p><strong>Variation cannot simply be ignored either.</strong> The fifth run really was wrong.
“Agents are stochastic” is not an excuse to treat every output as acceptable,
and a design that merely tolerates difference would have passed it.</p>

<p>The useful distinction was not whether the trajectories matched. It was whether
the resulting work satisfied the task contract.</p>

<blockquote>
  <p>Behavioral fingerprints can describe <em>how</em> an agent changed, but task outcomes
must decide <em>whether</em> the change is a regression.</p>
</blockquote>

<p>Trace similarity may still be valuable. A sudden increase in repeated tool
calls, malformed actions, recovery loops, or unnecessary steps can help explain
a failure. But these measurements should not silently become release gates
unless the user has explicitly declared them part of the product contract.</p>

<p>Otherwise an observability signal becomes an accidental definition of
correctness.</p>

<h2 id="sometimes-the-comparison-itself-is-invalid">Sometimes the comparison itself is invalid</h2>

<p>A second experiment exposed a different failure mode.</p>

<p>I attempted to continue archived agent trajectories on a new GPU host. The model
identifier, model revision, vLLM version, CUDA version, prompt, and scaffold
were held constant. The original trajectories had been generated on an A100
system; the new host used an H100.</p>

<p>Before running the full continuation experiment, I registered a serving-stack
behavioral compatibility check. At each archived fork point, the new host was
asked to generate the next action from the frozen prefix.</p>

<p>There were 24 distinct fork roots. One could not be compared because the
original malformed response had not been preserved, leaving 23
archived-comparable roots.</p>

<p>The result was unambiguous:</p>

<ul>
  <li>Raw responses matched on <strong>0 of 23</strong> roots.</li>
  <li>Structured actions matched on <strong>0 of 23</strong> roots.</li>
</ul>

<p>In some cases both systems performed broadly similar exploration using different
commands. In others the difference was categorical — the archived agent was
writing a patch while the new host was still reading source code:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>B_0 @ h=16   archived  find /testbed -path "*test*" -name "*.py" -exec grep -l "caplog" ...
             new host  find /testbed -path "*testing*" -name "*.py" -type f ...

A_3 @ h=24   archived  cat &gt; /tmp/fix_clear2.py &lt;&lt; 'EOF' ...
             new host  sed -n '685,700p' /testbed/src/_pytest/logging.py
</code></pre></div></div>

<p>It would have been easy to call this a behavioral regression. That would have
been wrong. <strong>The agent revision had not changed. The serving environment had.</strong></p>

<p>The appropriate result was:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>INCOMPARABLE — serving fingerprint changed
</code></pre></div></div>

<p>This is more than a reproducibility footnote. Without a comparability state, an
evaluation system is forced to misclassify environmental drift as agent
behavior.</p>

<p>A regression framework for agents therefore needs at least three possibilities:</p>

<ul>
  <li>the candidate is comparable and did not regress;</li>
  <li>the candidate is comparable and did regress;</li>
  <li>the candidate is not comparable to the baseline.</li>
</ul>

<p>Treating the third case as a failed test hides the actual cause and encourages
teams to debug the agent when the measurement instrument has moved.</p>

<p>One caveat I want to keep attached to that number: this is <strong>one</strong> stack change
on <strong>one</strong> workload. The GPU, the driver and the kernel path moved together, so
nothing here isolates a cause, and <code class="language-plaintext highlighter-rouge">temperature 0</code> is not token-level
determinism across execution stacks. What it supports is narrow and sufficient:
a serving-stack change is not safely treated as an ordinary candidate mutation.</p>

<h2 id="a-better-decision-order">A better decision order</h2>

<p>These observations led to the current AgentSeism decision flow.</p>

<figure>
  <img src="/assets/img/agentseism-decision-order.png" alt="A three-step decision flow. Step one, the comparability check, exits to INCOMPARABLE when the serving fingerprint changed, having spent zero candidate trials; otherwise it passes to step two. Step two, the outcome gate, exits to PASS_WITH_CHANGE when no gate is crossed and the trajectory moved but the outcome held, to INSUFFICIENT_EVIDENCE when there is too little valid data — which is not a pass — and to REGRESSION when a gate is crossed. Only REGRESSION continues to step three, trace root-cause analysis, which produces a localised suspect stage." loading="lazy" />
  <figcaption><b>Figure 2.</b> Each step can end the check. Only a confirmed outcome regression reaches
  root-cause analysis, and the cheapest check runs first.</figcaption>
</figure>

<h3 id="1-check-comparability-before-running-trials">1. Check comparability before running trials</h3>

<p>The framework first compares a serving fingerprint containing the execution
properties the user has declared invariant: model identifier and revision,
runtime and dependency versions, GPU and driver, decoding and context
parameters, prompt and scaffold versions, task image digest, evaluator version.</p>

<p>If a required field changes, AgentSeism returns <code class="language-plaintext highlighter-rouge">INCOMPARABLE</code> <strong>before spending
money on candidate trials</strong>.</p>

<p>This check must happen first, and the reason is not only tidiness. Running a
full evaluation and discovering afterward that the comparison was invalid
converts an inexpensive metadata check into a wasted experiment. Detected first
it costs zero trials; detected last it costs the entire budget for a comparison
nobody can use.</p>

<h3 id="2-ground-the-verdict-in-outcomes">2. Ground the verdict in outcomes</h3>

<p>For comparable evaluations, the release verdict comes from features defined in a
user contract: task success, evaluator-resolved status, safety violations,
latency or cost limits, tool reliability, completion under a declared step
budget.</p>

<p>Each feature has an explicit role. Some are release gates; others are warnings
or descriptive measurements. <strong>The framework should not decide after seeing the
results which metric mattered</strong> — which is why the contract is declared up front
and hashed into every report.</p>

<p>There is a usability tension here worth naming. A developer should not need to
understand risk difference or paired bootstrap to run a check, but a metric must
not be chosen after the analysis. AgentSeism resolves it by expanding a
three-line contract against versioned defaults into the complete contract that
is actually validated, hashed, and printed. Defaults are fixed in a released
version before anyone sees a number, so a default chosen in advance is not a
metric chosen afterwards — and the report shows every field, including the ones
the author never typed.</p>

<p>Repeated trials are sampled and analyzed at the task level rather than
pretending that multiple runs of one scenario are independent tasks.</p>

<h3 id="3-run-trace-rca-only-after-a-regression">3. Run trace RCA only after a regression</h3>

<p>If the outcome gate detects a regression, traces become useful diagnostic
evidence. Did failures begin after a tool error? Did the agent exhaust its step
budget? Did recovery behavior change? Did a particular stage accumulate extra
retries? Did the candidate stop validating its work?</p>

<p>This is root-cause analysis, not verdict authority.</p>

<p>Keeping RCA off the main decision path has a practical effect: a
suspicious-looking trace cannot turn a passing outcome into a regression unless
the user explicitly placed that trace feature in the contract. In the
implementation the renderer refuses to print RCA output under any verdict other
than <code class="language-plaintext highlighter-rouge">REGRESSION</code>, because a reader would take it as a reason the merge is
risky.</p>

<h2 id="four-verdicts-not-one-score">Four verdicts, not one score</h2>

<table>
  <thead>
    <tr>
      <th>Verdict</th>
      <th>Meaning</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">PASS</code> / <code class="language-plaintext highlighter-rouge">PASS_WITH_CHANGE</code></td>
      <td>Comparable, and no registered outcome regression. Behavioral change may still be reported.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">REGRESSION</code></td>
      <td>A registered outcome feature crossed its threshold under comparable conditions.</td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">INSUFFICIENT_EVIDENCE</code></td>
      <td>Not enough valid independent evidence. <strong>This is not a pass.</strong></td>
    </tr>
    <tr>
      <td><code class="language-plaintext highlighter-rouge">INCOMPARABLE</code></td>
      <td>A required execution or serving condition changed, so a behavioral comparison would mislead.</td>
    </tr>
  </tbody>
</table>

<p>The distinction between <code class="language-plaintext highlighter-rouge">PASS_WITH_CHANGE</code> and <code class="language-plaintext highlighter-rouge">REGRESSION</code> is central. It lets
a report say: <em>the agent behaved differently, but the registered product outcome
did not degrade.</em></p>

<p>The distinction between <code class="language-plaintext highlighter-rouge">INSUFFICIENT_EVIDENCE</code> and <code class="language-plaintext highlighter-rouge">PASS</code> is equally important,
and easier to lose. A small or invalid sample should not become approval merely
because no statistically detectable failure appeared. In practice this row is
the one most likely to be read as safe, so the report states it in words:
<em>neither safe nor unsafe — do not read this as a pass.</em></p>

<h2 id="what-agentseism-is-and-is-not">What AgentSeism is, and is not</h2>

<p>AgentSeism sits between an agent evaluation harness and CI. The harness executes
tasks and records outcomes. AgentSeism applies a predeclared contract, checks
comparability, compares baseline and candidate, and produces a report suitable
for a pull request.</p>

<p>It is not trying to replace domain-specific evaluators. A coding agent still
needs tests or a task grader; a support agent may need policy and resolution
checks; a security agent may need deterministic validation of its actions.</p>

<p>It is also not a generic “agent score”. Compressing correctness, cost,
reliability, and trace behavior into one number would erase the distinction the
system exists to preserve — and a single number needs weights, which is how a
severe recovery regression gets averaged away by three healthy metrics.</p>

<p>The intended question is narrower:</p>

<blockquote>
  <p>Given this declared contract, is the candidate comparable to the baseline, and
is there enough outcome evidence to call the change a regression?</p>
</blockquote>

<h2 id="the-next-test">The next test</h2>

<p>The current evidence establishes two concrete failure modes:</p>

<ol>
  <li>correct agent runs can follow substantially different trajectories;</li>
  <li>a serving-stack change can make apparently matched evaluations behaviorally
incomparable.</li>
</ol>

<p>It does not yet establish how well the complete workflow performs as a CI
method.</p>

<p>To test that, I have preregistered a small pilot using controlled engineering
mutations. It compares a baseline against a reduced agent step limit, and
against a weakened recovery hint following a controlled malformed-action
challenge. Three held-out tasks, three arms, two repetitions per cell —
18 runs. Its purpose is descriptive feasibility, not a general claim about all
agents or tasks.</p>

<p>The experiment was registered before execution: mutation values, task-selection
rules, run order and its hash, timeout handling, cost limits, and the
interpretations allowed for negative or insufficient results. Two of those
deserve mention because they are the parts most easily bent afterwards.</p>

<p><strong>Obtaining a regression is not a success condition.</strong> If a registered mutation
produces no regression, the null result stands. The mutation is not
strengthened, the tasks are not changed, the threshold is not lowered, and
repetitions are not added to earn one.</p>

<p><strong>A timeout is not a task failure.</strong> The pilot has a per-run wall-clock cap for
cost control. A run that hits it is censored, not failed — otherwise a budget
rule would manufacture a regression out of long tasks, and long tasks are
exactly what a reduced step limit is hypothesised to affect.</p>

<p>At the time of writing, the real pilot has not been run. That distinction is
deliberate. The framework should not manufacture a regression merely to complete
a demo, and this article should not turn an unobserved result into a product
claim.</p>

<h2 id="why-this-matters">Why this matters</h2>

<p>Stochastic systems need regression testing, but they cannot inherit
deterministic testing assumptions unchanged.</p>

<p>Gate on exact behavior, and you reject harmless variation. Ignore behavior
entirely, and you miss real failures. Compare runs produced by different serving
conditions without checking compatibility, and you attribute to the agent what
came from the measurement environment.</p>

<p>A useful Agent CI system therefore needs three separate concepts:</p>

<ul>
  <li><strong>Comparability</strong> — are these runs valid counterparts?</li>
  <li><strong>Outcome regression</strong> — did a user-important result degrade?</li>
  <li><strong>Diagnosis</strong> — where in the trajectory did the degradation likely arise?</li>
</ul>

<p>The ordering is the method:</p>

<blockquote>
  <p>Comparability first. Outcomes decide. Traces explain.</p>
</blockquote>

<p>AgentSeism is an attempt to make that ordering executable rather than leaving it
as evaluation advice.</p>

<p>The project is still early. The most valuable next evidence will not be another
abstract metric. It will be whether developers can declare a small contract, run
a real agent change, and use the resulting report to make a merge decision.</p>

<p>That is the standard the project now has to meet.</p>]]></content><author><name>Ruxi Zhang</name></author><summary type="html"><![CDATA[Regression testing assumes repeated runs are stable. Agents break that assumption. Two experiments — four correct runs that shared no trajectory, and 23 of 23 fork points diverging when only the GPU changed — pushed AgentSeism to a different order: comparability first, outcomes decide, traces explain.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://lyr-ai.github.io/assets/img/agentseism-two-failure-modes.png" /><media:content medium="image" url="https://lyr-ai.github.io/assets/img/agentseism-two-failure-modes.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Memory is not a list of facts</title><link href="https://lyr-ai.github.io/memory-is-not-a-list-of-facts/" rel="alternate" type="text/html" title="Memory is not a list of facts" /><published>2026-09-13T10:00:00-07:00</published><updated>2026-09-13T10:00:00-07:00</updated><id>https://lyr-ai.github.io/memory-is-not-a-list-of-facts</id><content type="html" xml:base="https://lyr-ai.github.io/memory-is-not-a-list-of-facts/"><![CDATA[<figure>
  <img src="/assets/img/memory-as-state-hero.png" alt="One memory on two time axes. On the valid-time axis, 'lives in San Jose' runs from two years ago to a month ago and 'lives in Seattle' from a month ago onward. On the observed-time axis, the user's original statement two years ago, a model inference last week, and the user's 'I moved a month ago' yesterday. The inference, confidence 0.95 but authority 0.3, bounces off a provenance guard; the user's statement, observed yesterday, sets validity from a month ago." loading="lazy" />
  <figcaption><b>Figure 1.</b> Same memory, two axes. The inference is more confident and more recent than the statement it would overwrite — and loses anyway.</figcaption>
</figure>

<p><em>What building long-term memory for agents taught me about state.</em></p>

<p>When I started building long-term memory for agents, I thought the hard
problem was retrieval. Store memories, embed them, retrieve the relevant ones,
rank them well enough that the agent sees the right context.</p>

<p>That model works — until memory starts changing.</p>

<p>Two years ago, a user said: <em>“I live in San Jose.”</em> Last week, the model
inferred from a conversation: <em>“The user lives in Seattle.”</em> Which one should
the system believe?</p>

<p>At first this looks like a retrieval question: retrieve both, let the model
decide. But the moment an agent is expected to hold persistent state across
months, it stops being about retrieval. It is a state-management problem, and
it comes with the questions state-management problems always come with. Who
is allowed to change this? When was it true, as opposed to when did we learn
it? Can we reconstruct how we got here? Does every kind of state change the
same way?</p>

<p>I ran into each of those while building <a href="https://github.com/lyr-ai/typedmem">TypedMem</a>.
Here is what each one changed.</p>

<h2 id="1-confidence-is-not-authority">1. Confidence is not authority</h2>

<p>The San Jose memory came from an explicit user statement. The Seattle memory
came from a model inference, with high confidence, made last week.</p>

<p>A policy that ranks by recency and confidence prefers Seattle. That is
uncomfortable, and the discomfort is the point: the two memories differ on a
dimension neither recency nor confidence measures.</p>

<p>Confidence answers <em>how sure are we about this claim</em>. Authority answers <em>how
entitled is this source to override another</em>. An explicit statement from the
user and an inference by the model are not interchangeable evidence, however
sure the model is.</p>

<p>So authority became its own guard: a memory with weaker provenance cannot
replace one with stronger provenance. Not folded into confidence, not weighted
into a score — checked first, on its own.</p>

<p>The reverse is deliberately not true. Higher authority does not automatically
win. An old high-authority statement should not overwrite a newer state just
because its source is stronger; the user may well have moved. Authority is a
veto for the weaker side, not a trump card for the stronger one.</p>

<p>The tempting alternative was a single number:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>score = a·authority + b·confidence + c·recency
</code></pre></div></div>

<p>It is attractive because it produces one ordering. It also destroys the
meaning of each term. Authority, confidence and time are different kinds of
evidence, and a conflict rule that keeps them separate is slightly more
explicit and much easier to reason about — you can say <em>why</em> a memory lost.</p>

<h2 id="2-observation-time-is-not-validity-time">2. Observation time is not validity time</h2>

<p>The original schema had one timestamp, and it was quietly doing three jobs:
when the memory was observed, when its content became true, and where
confidence decay starts. Those coincide often enough that one field seems
fine.</p>

<p>Then: yesterday, the user says <em>“I moved to Seattle a month ago.”</em></p>

<p>There are two times now. The memory was observed yesterday; its content has
been true for a month. And the previous state has a natural end: San Jose
stopped being true a month ago, not yesterday.</p>

<p>So a memory carries a validity window, half-open — <code class="language-plaintext highlighter-rouge">[valid_from, valid_to)</code> —
separate from the observation time. Half-open so that when one state ends and
the next begins at the same instant, exactly one of them is valid.</p>

<p>Two consequences I did not anticipate.</p>

<p><strong>Future state becomes expressible.</strong> <em>“Starting October 1, I’ll be using
Postgres”</em> is observed today and valid later. Which forces a rule: filter by
validity <em>first</em>, then pick the latest — otherwise a declared future state
shadows the current one merely because its start is later.</p>

<p><strong>“Unspecified” has to stay unspecified.</strong> For old data with no validity
window, the observation time is the operational fallback. But the storage
keeps the distinction: a missing <code class="language-plaintext highlighter-rouge">valid_from</code> means <em>the writer did not say</em>,
not <em>the writer said it starts at observation</em>. That sounds pedantic until
you need to replay history and cannot tell which memories were declared and
which were defaulted.</p>

<h2 id="3-history-is-not-replay">3. History is not replay</h2>

<p>The system already had an event log. A replacement produced an event saying
one memory replaced another, with the new version number and a fragment of
the old content. Good for auditing. Useless for reconstruction.</p>

<p>The gap is between <em>what happened</em> and <em>what the state was</em>. Knowing that
memory 42 was replaced at version 3 does not tell you what memory 42 looked
like before, or exactly what it became. A log of events is a history of
actions; it is not a history of state.</p>

<p>The fix was less clever than I expected. Every state-changing event now
carries the full logical state of the memory before and after. Not a diff —
the whole thing, twice. Creation has no <em>before</em>; deletion has no <em>after</em>.</p>

<p>That costs storage, and it makes replay boring: walk the events in order,
keep the <em>after</em>. No patch language, no dependency on an earlier snapshot, no
ambiguity about nested structure. Boring is the property I wanted.</p>

<p>One rule matters more than the format:</p>

<blockquote>
  <p>Replay must not re-run conflict resolution.</p>
</blockquote>

<p>Suppose the conflict policy changes six months from now. If replay feeds the
historical inputs through today’s policy, the reconstructed past differs from
the past the system actually produced. Replay restores decisions; it does not
reconsider them. The same log, under any current policy, replays to the same
historical state.</p>

<p>That distinction — <em>restore what was decided</em> versus <em>re-decide from the
inputs</em> — turned out to be the same one I keep meeting in agent runtimes:
replaying a recorded execution and re-executing it are different products,
and a system that blurs them is wrong about both.</p>

<h2 id="4-there-may-be-no-universal-rule-for-who-wins">4. There may be no universal rule for who wins</h2>

<p>Once authority and validity were explicit, one more hard-coded assumption
became visible. Under replacement, an incoming memory had to be no weaker than
the existing one on <em>both</em> validity start and confidence.</p>

<p>The tempting abstraction was a lexicographic ordering: compare validity start
first, then confidence as a tie-breaker. It is a natural thing to write, and
it quietly changes the semantics. Timestamps rarely tie, so confidence would
almost never be consulted. A newer memory with confidence 0.5 would replace a
day-old one with confidence 0.9. That is <em>newest wins</em> with extra steps.</p>

<p>What the rule actually was — and what I kept — is a <strong>conjunctive guard</strong>: to
replace, the incoming memory must be no weaker on <em>every</em> listed dimension.
Newer and stronger replaces; newer but weaker is ignored; older is ignored.
The order of the dimensions does not matter, because nothing is being broken
in sequence — every one is checked.</p>

<p>And then the question that made this a per-type decision: should every kind of
memory use the same guard? A deadline may care primarily about the latest
effective state. A biographical fact may require both recency and confidence.
Other memory types will make different trade-offs. These are different kinds
of state, and there is no reason they share update semantics. So which
dimensions guard replacement is declared by the memory type.</p>

<p>With one exception. In TypedMem, authority stays outside that list: a type may
choose which temporal and confidence signals guard replacement, but it cannot
silently disable provenance protection. Letting a configuration quietly turn
off a safety invariant would turn the first lesson back into a suggestion.</p>

<h2 id="what-changed-in-my-mental-model">What changed in my mental model</h2>

<p>None of these are four features. They are consequences of one decision:
treating memory as state rather than as a collection of documents.</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>retrieval asks       what is relevant to this query?
memory-as-state asks what do I currently believe, why, since when,
                     who was allowed to change it, and can I reconstruct
                     how I got here?
</code></pre></div></div>

<p>The system has more concepts now than it did — authority, validity, before
and after, per-type guards. On paper that is more machinery. But the
underlying problems existed before the concepts did. A single timestamp did
not remove temporal semantics; it forced three kinds of time into one field.
Ignoring authority did not remove provenance; it let confidence make decisions
it was never designed to make. An audit log without state did not remove the
need for replay; it made replay impossible. The distinction that matters is
not simple versus complex. It is essential versus accidental complexity, and
the test is whether each concept means exactly one thing.</p>

<h2 id="what-i-deliberately-did-not-build">What I deliberately did not build</h2>

<p>Automatic closing of the previous validity window. Ranking by authority.
Compaction and checkpoints for the event log. Migrating history under a new
policy. Rebuilding a store from its log. A general rule language for
resolution.</p>

<p>Some of these will turn out to be needed. None of them has yet, and an
abstraction looks most attractive right before there is evidence for it. So
the contract stops here: memories carry provenance, confidence, observation
time and validity; types declare how replacement is guarded; every mutation
leaves enough behind to reconstruct its result.</p>

<p>The next step is not another rule. It is measuring whether these semantics
actually make agents more reliable. That is a question for a benchmark, not
another design note.</p>]]></content><author><name>Ruxi Zhang</name></author><summary type="html"><![CDATA[Four things building long-term agent memory taught me about state: confidence is not authority, observation time is not validity time, an audit log is not a replayable history, and there may be no universal rule for which memory wins.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://lyr-ai.github.io/assets/img/memory-as-state-hero.png" /><media:content medium="image" url="https://lyr-ai.github.io/assets/img/memory-as-state-hero.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">Restoring the code is not resuming the agent</title><link href="https://lyr-ai.github.io/restoring-the-code-is-not-resuming-the-agent/" rel="alternate" type="text/html" title="Restoring the code is not resuming the agent" /><published>2026-09-12T14:00:00-07:00</published><updated>2026-09-12T14:00:00-07:00</updated><id>https://lyr-ai.github.io/restoring-the-code-is-not-resuming-the-agent</id><content type="html" xml:base="https://lyr-ai.github.io/restoring-the-code-is-not-resuming-the-agent/"><![CDATA[<p><em>What agent replay taught me about durable execution.</em></p>

<p>While building replay support for
<a href="https://github.com/lyr-ai/agentseism">AgentSeism</a>, I tried to resume two
coding-agent executions from exactly the same repository state. The source
fingerprint matched byte for byte. The executions did not.</p>

<p>One carried a different scratch file. Their message histories were different.
Restoring the repository gave me a state that looked correct according to my
instrumentation, but it was not the same agent execution.</p>

<p>That failure changed how I think about long-running agents. Once an agent
executes tools, modifies an environment, accumulates context, and runs for
minutes or hours — on one fixed task I measured runs from three minutes to
over six hours — recovery stops looking like “reload the prompt and workspace”
and starts looking like a distributed-systems problem.</p>

<p>I had started with a simple question: “How do I save the workspace?” The
experiment forced a harder one: <strong>“What does it actually mean to resume an
agent?”</strong> Following it led quickly into checkpoint consistency, execution
ownership, fencing, idempotency, replay semantics, and external side effects.</p>

<h2 id="1-repository-state-is-only-part-of-agent-state">1. Repository state is only part of agent state</h2>

<p>For a coding agent, the repository is an obvious piece of execution state and
one of the easiest pieces to measure objectively. In AgentSeism, tracked source
state is represented using a canonical Git diff, so two executions can be
compared without interpreting model reasoning.</p>

<p>That is what made the pair in the opening measurable in the first place: the
same canonical diff, byte for byte. Everything the diff does not cover —
scratch files, message history — was different. The source code was identical;
the complete execution state was not.</p>

<figure class="wide">
  <img src="/assets/img/agent-execution-state.png" alt="Agent execution state decomposed into message history, tracked repository state, untracked scratch workspace, tool observations, step/token/cost budgets, and external-effect state." loading="lazy" />
  <figcaption><b>Figure 1.</b> The repository is one component of agent execution state, not the state itself.</figcaption>
</figure>

<p>The repository is therefore one component of agent state rather than the agent
state itself. <strong>Restoring repository state is not the same as resuming the
agent execution.</strong></p>

<p>Restore the transcript and the tracked source but drop the scratch files, and
you get an agent whose memory contradicts what it can see: it remembers writing
<code class="language-plaintext highlighter-rouge">repro.py</code>, reads it back, and is told the file does not exist. What follows is
not a resumed job. It is a confused one.</p>

<h2 id="2-a-checkpoint-is-a-recoverable-boundary">2. A checkpoint is a recoverable boundary</h2>

<p>A more useful definition is: <strong>a checkpoint is a recoverable boundary in the
execution state machine.</strong></p>

<p>For a long-running coding agent, that boundary may include conversation
context, tracked and untracked workspace, tool state, execution step, remaining
token/time/cost budgets, model configuration, and progress of externally
visible operations. The exact schema is application-dependent. The important
property is that the runtime can give a precise meaning to recovery from that
checkpoint.</p>

<p>Restarting a fresh agent from an old Git diff may be useful, but it is a
restart from source state, not a resume of the original execution. The two are
different products, and a runtime should say which one it offers.</p>

<h2 id="3-partial-checkpoints-are-more-dangerous-than-missing-checkpoints">3. Partial checkpoints are more dangerous than missing checkpoints</h2>

<p>Suppose an agent begins creating a checkpoint at step 100. Messages and
workspace are stored, but the worker crashes before tool state is durable.
Restoring the pieces that happen to exist would create a mixed state that never
existed in the original execution.</p>

<p>This is worse than having no checkpoint at all, because it looks legitimate. A
missing checkpoint costs work; a half-written one costs trust, and the failure
surfaces much later as agent behaviour nobody can explain.</p>

<p>A safer design writes checkpoint components as immutable objects first and
publishes a small manifest only after every required component is durable.</p>

<figure class="wide">
  <img src="/assets/img/data-first-pointer-last.png" alt="An agent at step 100 writes messages, workspace and tool-state blobs to an object store; only when all required objects are durable is a checkpoint manifest atomically published as a valid recovery point, otherwise nothing is published." loading="lazy" />
  <figcaption><b>Figure 2.</b> Data first, pointer last. Recovery honours committed manifests only; loose blobs are never a checkpoint.</figcaption>
</figure>

<p>Recovery recognizes only a committed manifest. Partially uploaded objects are
harmless orphans until garbage collection or later deduplication. The goal is
to make invalid mixed states impossible to express, rather than merely
detectable after the fact.</p>

<p>This is not two-phase commit. It is the much cheaper pattern — <em>data first,
pointer last</em> — that Git uses for objects and refs, and that filesystem
journals and object-store table formats use for the same reason: make the only
mutable thing a single small atomic write.</p>

<h2 id="4-immutable-data-and-execution-authority-are-different-problems">4. Immutable data and execution authority are different problems</h2>

<p>Suppose Worker A owns an execution under generation 7. A network partition
causes its lease to expire, and Worker B takes ownership under generation 8.
Worker A may still be alive.</p>

<p>If A finishes uploading an immutable, content-addressed workspace blob, little
harm occurs. What A must not be allowed to do is declare a new recovery point
after it has lost ownership.</p>

<p>This leads to a useful principle: <strong>separate immutable data-plane writes from
ownership-sensitive metadata publication.</strong> Blob writes can be independent;
publishing a checkpoint manifest or advancing the recovery pointer requires the
current fencing generation.</p>

<h2 id="5-the-architecture-naturally-splits-into-a-control-plane-and-a-data-plane">5. The architecture naturally splits into a control plane and a data plane</h2>

<p>As the recovery requirements accumulated, the control-plane/data-plane boundary
became the architectural abstraction I found most useful.</p>

<p>The <strong>control plane</strong> owns execution intent and authority: job lifecycle,
scheduling, leases, fencing generations, recovery lineage, quota, policy, and
checkpoint publication authority.</p>

<p>The <strong>data plane</strong> performs the work: workers, sandboxes, the agent loop, model
and tool calls, workspace mutation, immutable checkpoint production, and trace
emission.</p>

<figure>
  <img src="/assets/img/control-plane-data-plane.svg" alt="Control plane (job API, durable queue, scheduler, lease/fencing, checkpoint metadata, recovery head) above a data plane (worker, sandbox, agent runtime, tool/effect service, model gateway). Three edges cross the boundary: the scheduler issues a lease with a generation to the worker; the agent runtime writes immutable blobs to the checkpoint blob store from any generation; and the agent runtime publishes a manifest to checkpoint metadata, fenced by generation." loading="lazy" />
  <figcaption><b>Figure 3.</b> Control plane and data plane for a durable agent runtime. The blob write is unconditional; the manifest publish is a request the control plane may refuse.</figcaption>
</figure>

<p>Two edges cross the boundary upward, and they are different in kind. The blob
write is unconditional: content-addressed, immutable, harmless from any
generation. The manifest publish is a request that the control plane may
refuse, and refuses whenever the generation is stale.</p>

<p>The distinction is not merely organizational. The data plane is where risky and
failure-prone execution happens. A buggy agent or crashed worker should not be
able to redefine execution ownership or declare an arbitrary checkpoint
authoritative.</p>

<p><strong>The data plane may produce candidate state; the control plane decides which
state is authoritative.</strong></p>

<h2 id="6-lease-fencing-and-recovery-preserve-authority">6. Lease, fencing, and recovery preserve authority</h2>

<p>A lease defines time-bounded execution ownership, but it cannot prove that an
old worker has stopped. During a network partition, a stale worker may continue
running even after the control plane has reassigned the job. For a while, two
workers are physically executing the same job.</p>

<p>Fencing therefore protects ownership-sensitive writes. Each assignment receives
a monotonically increasing generation, and durable services reject mutations
from stale generations. The store does the rejecting, not the worker, because a
partitioned worker cannot be trusted to know it has been fenced.</p>

<p>Worker-side self-termination is still useful because it reduces wasted
computation, but it is not the correctness boundary. <strong>Self-fencing is an
optimization; server-side fencing is the correctness boundary.</strong></p>

<p>The four mechanisms solve different problems, and each catches what the one
above it cannot:</p>

<table>
  <thead>
    <tr>
      <th>mechanism</th>
      <th>catches</th>
      <th>cannot catch</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>lease</td>
      <td>a dead worker</td>
      <td>a live, partitioned one</td>
    </tr>
    <tr>
      <td>fencing</td>
      <td>its writes to our stores</td>
      <td>what it already sent outside</td>
    </tr>
    <tr>
      <td>idempotency</td>
      <td>duplicate external effects</td>
      <td>providers that cannot deduplicate</td>
    </tr>
    <tr>
      <td>self-fencing</td>
      <td>continued waste</td>
      <td>anything, as a correctness claim</td>
    </tr>
  </tbody>
</table>

<p>The bottom row is the trap: self-fencing is the only one of the four that is
not a correctness mechanism, and it is the one that looks most like a fix.</p>

<h2 id="7-latest-is-not-one-thing">7. “Latest” is not one thing</h2>

<p>Suppose an agent reaches step 120 but its newest durable checkpoint is step</p>
<ol>
  <li>The worker crashes, and a replacement resumes from step 100 under a higher
generation. Ownership moved forward while logical execution progress moved
backward. Both statements are correct, and a design with one sequence number
has to lie about one of them.</li>
</ol>

<p>A runtime should therefore distinguish three orderings, and use each for what
it is for:</p>

<ul>
  <li><strong>generation</strong> — who owns execution now. Strictly monotonic. Recovery
selection uses this first.</li>
  <li><strong>checkpoint sequence</strong> — which checkpoint within this attempt. Monotonic
within a generation. Recovery selection uses this second.</li>
  <li><strong>agent step</strong> — how far this branch has got. <em>Not</em> monotonic across
recovery. Billing and budgets use it; recovery selection never does.</li>
</ul>

<p>Choosing a recovery point by agent step is how a stale branch gets adopted
because it happened to get further. I prefer <strong>recovery_head</strong> over
<code class="language-plaintext highlighter-rouge">latest_checkpoint</code>: it means the newest durable recovery point on the
currently authoritative execution lineage, not the largest historical step
number.</p>

<p>The user-facing consequence is small but real: a progress indicator derived
from the recovery head goes <code class="language-plaintext highlighter-rouge">120 → 100</code> after a recovery, which reads as data
loss for a system that behaved correctly. Report the high-water mark and the
current branch position as two numbers, and do not invent a percentage — an
agent’s step count is not distance-to-completion.</p>

<h2 id="8-checkpointing-does-not-make-external-effects-transactional">8. Checkpointing does not make external effects transactional</h2>

<p>The control plane can protect state it owns. The external world is harder.</p>

<p>Suppose an agent checkpoints, calls <code class="language-plaintext highlighter-rouge">send_email()</code>, the provider accepts the
email, and the worker crashes before recording success. A replacement restores
the previous checkpoint. The runtime cannot know whether retrying will
duplicate the message.</p>

<p>No better workspace snapshot removes this ambiguity. A durable runtime
therefore needs a separate effect protocol: write-ahead intent, durable logical
operation identity, effect execution, result commitment, and reconciliation
when the outcome is uncertain.</p>

<p>The write-ahead intent is what turns an unknown into a known unknown. On
recovery, <code class="language-plaintext highlighter-rouge">pending</code> means <em>this may or may not have happened</em> — not a solution,
but the precondition for every one that follows. The cheapest discharge is
often to trade readability for idempotency: embed the operation id in the
message itself, and on recovery search the provider’s sent items before
deciding whether to send. Where no read path exists, the tool declares whether
it prefers at-least-once or at-most-once, and for the latter the platform parks
the job for a human rather than guessing.</p>

<p><strong>Checkpointing recovers computation state; it does not make external side
effects transactional.</strong></p>

<h2 id="9-recovery-and-replay-can-want-different-observations">9. Recovery and replay can want different observations</h2>

<p>Read-only tools avoid duplicate effects but introduce another problem. Suppose
an expensive search returns observation O1 after the last checkpoint, and the
worker later crashes. Re-running the search may cost money and may now return
O2.</p>

<p>Replay asks, “What did the agent observe then?” and usually wants O1. Recovery
asks, “What should the live job observe now?” and may prefer O2.</p>

<figure>
  <img src="/assets/img/recovery-vs-replay.png" alt="An original observation O1 recorded at time t0 feeds two paths: replay, which reproduces historical execution and prefers the recorded O1; and recovery, which continues live execution and asks whether freshness is semantically important — if not, reuse O1; if so, re-observe O2 and record the divergence." loading="lazy" />
  <figcaption><b>Figure 4.</b> Replay and recovery can require different semantics for the same observation.</figcaption>
</figure>

<p>For expensive observations, persistence can be decoupled from checkpoint
cadence: record the result at the moment of the call, keyed by its semantic
input, so that a crash does not force a high checkpoint frequency on a system
that did not otherwise need one. Semantic cache identity must include all
result-changing context, including tool version, canonical arguments,
repository revision, and tenant or authorization scope. Omitting authorization
scope can turn a performance optimization into a cross-tenant data leak.</p>

<p>Retention is also part of the contract. Recovery-critical payloads cannot be
sampled away while recovery promises to use them. Observability payloads may
use sampling or tiered retention. Digests can be kept broadly to detect that an
observation changed even when the payload is no longer available.</p>

<p><strong>If recovery is expected to provide a payload, its retention period is a
correctness property, not merely a storage optimization.</strong></p>

<h2 id="10-resume-is-a-continuation-not-a-replay">10. Resume is a continuation, not a replay</h2>

<p>Everything above restores the starting point. None of it restores the future,
because the next model call is a fresh sample from a system that is not
deterministic.</p>

<p>I measured this directly. Forking continuations from an identically
reconstructed state — same tracked source, same scratch files, same
transcript, temperature 0:</p>

<div class="language-plaintext highlighter-rouge"><div class="highlight"><pre class="highlight"><code>14 of 16 continuations reproduced the donor's next three actions exactly
 0 of 15 reached the donor's final state
</code></pre></div></div>

<p>Short-horizon behaviour is near-deterministic; long-horizon behaviour is not.
The checkpoint guarantees the agent restarts from the same state, not that it
does the same thing.</p>

<p>The same applies to a retried model call. A response lost in transit and
re-requested is a new sample, not a repeat — even at temperature 0 the serving
stack does not reproduce itself, and in one batch of 333 model calls, 12 were
sampled more than once because of transport failures alone, one of them six
times. A silent retry does not redo the same work slightly wastefully; it
branches the trajectory, and the trajectory has to record that it did.</p>

<p>Two consequences follow. Idempotency keys cannot be derived from trajectory
position, because the resumed trajectory diverges from the original. And if a
product genuinely needs replay — audit, regression testing — it needs recorded
<em>outputs</em>, every tool result and every model response, replayed from the log
instead of re-executed. Replay from recorded outputs, yes; reproduce by
re-execution, no.</p>

<h2 id="11-long-running-agents-start-looking-like-durable-workflows">11. Long-running agents start looking like durable workflows</h2>

<p>A short agent can often be retried from the beginning. A two-hour agent cannot.
As runtime increases, so does accumulated context, completed tool work,
modified environment state, external effects, and compute already invested. The
probability of an infrastructure failure during the run also increases.</p>

<p>At that point, durability stops being optional. The system begins to resemble a
durable distributed workflow: a job has an owner, runs in an isolated worker,
advances an agent state machine, produces checkpoints and effects, and may
transfer ownership after failure.</p>

<p>This is why I expect agent infrastructure to borrow increasingly from workflow
engines, distributed schedulers, transactional systems, and sandboxed compute
platforms rather than only from model-serving APIs.</p>

<h2 id="12-what-i-would-build-next">12. What I would build next</h2>

<p>AgentSeism’s current replay and fork infrastructure is experimental rather than
a production runtime. Its purpose is to reconstruct execution states precisely
enough to study agent variation and intervention. But the work suggests a
concrete production direction.</p>

<p>I would make the control-plane/data-plane boundary explicit; make checkpoints
crash-consistent through immutable objects and atomic manifest publication;
route consequential external operations through a guarded effect service;
separate live recovery semantics from deterministic replay; and add stronger
sandbox isolation plus a shared inference layer.</p>

<p>I would then test the design with deliberate failure injection: worker crashes,
network partitions, stale owners, interrupted checkpoint publication, model
timeouts, and ambiguous tool effects.</p>

<p>The key metric would not be whether a resumed agent follows exactly the same
future trajectory. Section 10 says it will not. The stronger durability
question is whether the system can resume from a well-defined state without
corrupting execution, duplicating consequential effects, or losing information
required for correctness.</p>

<h2 id="conclusion">Conclusion</h2>

<p>I started with what looked like a small implementation task: reconstruct an
agent at a previous step so I could replay and fork its execution. The first
surprise was that restoring the repository was not enough.</p>

<p>Following that observation led naturally into checkpoint consistency, execution
ownership, fencing, idempotency, observation persistence, and recovery
semantics. None of these problems are unique to AI, but agents combine them in
an unusual way because they mix long-running computation, stochastic model
calls, mutable environments, and potentially irreversible tools.</p>

<p>The mental model I now use is simple:</p>

<p><strong>An agent checkpoint is not a saved workspace. It is a durable boundary in an
execution state machine.</strong></p>

<p>Once an agent becomes long-running enough to deserve recovery, that distinction
matters.</p>]]></content><author><name>Ruxi Zhang</name></author><summary type="html"><![CDATA[What agent replay taught me about durable execution: checkpoint consistency, execution ownership, fencing, and why resume is a continuation rather than a replay.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://lyr-ai.github.io/assets/img/recovery-vs-replay.png" /><media:content medium="image" url="https://lyr-ai.github.io/assets/img/recovery-vs-replay.png" xmlns:media="http://search.yahoo.com/mrss/" /></entry><entry><title type="html">One run is not an evaluation</title><link href="https://lyr-ai.github.io/one-run-is-not-an-evaluation/" rel="alternate" type="text/html" title="One run is not an evaluation" /><published>2026-09-11T17:00:00-07:00</published><updated>2026-09-11T17:00:00-07:00</updated><id>https://lyr-ai.github.io/one-run-is-not-an-evaluation</id><content type="html" xml:base="https://lyr-ai.github.io/one-run-is-not-an-evaluation/"><![CDATA[<figure>
  <img src="/assets/img/evaluation-hero.jpg" alt="run arrow score, struck through, above the words is not an evaluation; a chain reading protocol, repeat, classify, measure, label, uncertainty; and a list of five assumptions that broke." loading="lazy" />
  <figcaption><b>Figure 1.</b> Repeated runs reveal variance. Valid measurement tells you what varied. External outcomes tell you whether the variation mattered.</figcaption>
</figure>

<h2 id="abstract">Abstract</h2>

<p>I used to think evaluating an AI agent was conceptually straightforward: freeze
a task and configuration, run the agent, record the outcome, compare scores.</p>

<p>Repeated executions changed that view.</p>

<p>The same agent, task, model and <code class="language-plaintext highlighter-rouge">temperature=0</code> configuration can produce
materially different executions. More importantly, the evaluation system itself
can create artifacts that look like model behavior: a context-window limit can
selectively remove long trajectories, a reasonable-looking state definition can
create false reconvergence, an infrastructure timeout can masquerade as agent
failure, and an obviously “cold” benchmark run can already have been warmed by
an earlier failed measurement.</p>

<p>These are not implementation annoyances. They change what an evaluation result
means.</p>

<p>My current view is that agent evaluation needs at least five separate layers: a
frozen execution protocol, repeated runs, explicit termination and censoring
semantics, validated behavioral measurements, and an external task-level
outcome. Only after those are separated does it make sense to compare agents, or
to claim that a behavioral difference matters.</p>

<p>This post records the evaluation assumptions that broke while I was building
<a href="https://github.com/lyr-ai/agentseism">AgentSeism</a>, and the protocol I use now.</p>

<h2 id="1-one-run-represents-the-agent">1. “One run represents the agent”</h2>

<p>Suppose Agent A solves a benchmark task and Agent B fails it. Which agent is
better? With one execution each, we barely know.</p>

<p>A long-running agent is not a single model call. It repeatedly generates
actions, modifies an environment, observes the result, and folds those
observations into subsequent context, so small differences propagate through the
execution.</p>

<p>In my coding-agent experiments, repeated executions used the same task, model,
agent configuration and <code class="language-plaintext highlighter-rouge">temperature=0</code>, and still did not reliably follow the
same trajectory. Some differences disappeared, some persisted, and some
executions diverged, reached the exact same repository state, and diverged
again.</p>

<p>A single execution therefore measures something closer to</p>

\[Y_{i,r}\]

<p>— the outcome of run \(r\) on task \(i\) — than an intrinsic property of the agent.
What we usually care about is a distribution:</p>

\[P(Y \mid \text{agent}, \text{task}, \text{protocol}).\]

<p>Repeated execution is not simply a way to make an average more precise. It is
how you find out whether the thing being evaluated has meaningful run-to-run
variance at all.</p>

<figure class="wide">
  <img src="/assets/img/outcome-distribution.png" alt="A task with a frozen protocol fanning out into runs one through N, each producing an outcome, all feeding a single outcome distribution." loading="lazy" />
  <figcaption><b>Figure 2.</b> From a score to a distribution. Before comparing two agents, understand the variance of each under the protocol being used.</figcaption>
</figure>

<h2 id="2-a-failed-run-is-an-agent-failure">2. “A failed run is an agent failure”</h2>

<p>Repeated runs immediately raise another question: what counts as an outcome?</p>

<p>In one experiment I configured the serving system with a 32k context limit.
Several trajectories terminated with <code class="language-plaintext highlighter-rouge">ContextWindowExceeded</code>. One task — Seaborn
— lost all three of its runs that way.</p>

<p>It would have been easy to record <code class="language-plaintext highlighter-rouge">success: false</code> and fold those into the
agent’s accuracy. But the model itself supported a much longer context. The 32k
boundary was an infrastructure choice I had made. Those executions answered
“can this agent complete the task under my 32k serving constraint?” — not “can
this agent solve the task?”</p>

<p>The distinction got sharper when I noticed <em>which</em> executions were being
removed. Long trajectories involve more exploration and accumulate more context,
so context-window censoring was not random. The experiment was preferentially
deleting a particular kind of execution: a missing-not-at-random pattern.</p>

<p>I now separate at least four termination classes.</p>

<table>
  <thead>
    <tr>
      <th>Termination</th>
      <th>Interpretation</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Completed / submitted</td>
      <td>agent produced a terminal outcome</td>
    </tr>
    <tr>
      <td>Agent failure</td>
      <td>the agent itself terminated unsuccessfully</td>
    </tr>
    <tr>
      <td>Infrastructure failure</td>
      <td>endpoint, container, transport or runtime failed</td>
    </tr>
    <tr>
      <td>Censored</td>
      <td>an experimental boundary stopped the trajectory</td>
    </tr>
  </tbody>
</table>

<p>The exact taxonomy depends on the system. Collapsing these into <code class="language-plaintext highlighter-rouge">success=0</code> can
distort an evaluation badly.</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>                       EXECUTION ENDS
                             │
             ┌───────────────┼───────────────┐
             │               │               │
             ▼               ▼               ▼
         COMPLETED      AGENT FAILURE    NOT AN AGENT FAILURE
                                         ┌──────┴──────┐
                                         ▼             ▼
                                  INFRA FAILURE     CENSORED
                                  endpoint died     context limit
                                  container crash   step limit
                                  transport error   time boundary
</code></pre></div></div>

<p><strong>Figure 3.</strong> Failure is not one bucket. The question is not “did this run
finish?” but “what process caused observation to stop?”</p>

<h2 id="3-measurement-code-cannot-manufacture-a-finding">3. “Measurement code cannot manufacture a finding”</h2>

<p>The most uncomfortable evaluation bugs are not crashes. They are measurements
that produce plausible numbers. I hit three.</p>

<h3 id="a-cold-run-that-was-already-warm">A “cold” run that was already warm</h3>

<p>While measuring prefix caching, a script assumed <code class="language-plaintext highlighter-rouge">repeat == 0 → cold</code> and
<code class="language-plaintext highlighter-rouge">repeat == 1 → warm</code>. Reasonable enough. But earlier executions had already sent
the request and populated the cache before the script crashed while printing its
results. On rerun, the nominally first measurement was already warm.</p>

<p>The eventual data made it visible: a genuinely cold ~28.5k-token prefill took
about 12.1 seconds, while a cached repetition took about 0.43 seconds. Earlier
~14k measurements that had looked inexplicably fast were not evidence of
exceptional cold-prefill performance — the cache was already populated.</p>

<p>The bug was not in vLLM. It was in the assumption that <strong>the first recorded
measurement is the first system exposure.</strong> Experimental state can survive a
failed measurement.</p>

<h3 id="a-reconvergence-that-never-happened">A reconvergence that never happened</h3>

<p>AgentSeism compares trajectories by repository source state. The first
implementation treated two runs as reconverged if they shared a source-state
hash after diverging, and the analysis produced a lot of apparently anomalous
pairs.</p>

<p>The reason was embarrassing. The hash <code class="language-plaintext highlighter-rouge">e3b0c442…</code> is SHA-256 of the empty diff.
Every execution occupies that state before it modifies anything, so two
trajectories could “reconverge” purely by both having changed nothing yet. Once
the empty state was excluded, the anomalies largely disappeared.</p>

<p>Nothing about the agent changed. Only the measurement did.</p>

<h3 id="two-different-outcomes-that-were-the-same-source-state">Two “different outcomes” that were the same source state</h3>

<p>Another pair appeared to finish with different patches even though their final
tracked source hashes were identical. One submission contained roughly 18 KB of
extra untracked scratch content; the actual tracked modification was the same
~720-byte code change in both.</p>

<p>Two different objects had been compared: a <strong>submission string</strong> and a
<strong>canonical tracked repository state</strong>. Both are legitimate measurements of
different questions. If the question is whether the agent produced the same
source repair, the tracked state is the right outcome; if it is whether the
agent submitted the same artifact, the string may be. The bug was silently
treating the two as interchangeable.</p>

<figure class="wide">
  <img src="/assets/img/construct-to-result.png" alt="A chain from research question to construct, metric, instrumentation, validation and result." loading="lazy" />
  <figcaption><b>Figure 4.</b> The chain I now make explicit for any metric that carries weight in an argument.</figcaption>
</figure>

<h2 id="4-fixing-a-measurement-is-a-neutral-act">4. “Fixing a measurement is a neutral act”</h2>

<p>There is a dangerous moment in any exploratory experiment: the data look
strange, you find a flaw in the analysis, you fix it, the results look cleaner —
and then another strange result appears, and you fix that too. At some point
“debugging the measurement” becomes indistinguishable from adjusting the
analysis until the data tell a satisfying story.</p>

<p>I ran into exactly this.</p>

<p>The empty-state reconvergence issue was clearly a measurement bug: an unmodified
repository should not count as meaningful reconvergence. But another pattern
survived the fix — some trajectories reached the same <strong>non-empty</strong> source state
and later diverged again.</p>

<p>My original classification treated that as anomalous, because it implicitly
assumed that once trajectories reconverge they stay reconverged. The data showed
the assumption was false. Changing the instrumentation to make those cases
disappear would no longer have been fixing a measurement; it would have been
deleting a real topology my conceptual model failed to represent.</p>

<p>So I kept the observation and expanded the representation to a triple \((D, R, F)\) — divergence, non-empty reconvergence, final convergence — which describes
the three observed cases directly:</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>absorbed        1 1 1
persistent      1 0 0
re-divergence   1 1 0
</code></pre></div></div>

<blockquote>
  <p><strong>Fix instrumentation when it fails to measure the construct you defined.
Change the construct only when you are willing to say the original hypothesis
or representation was incomplete.</strong></p>
</blockquote>

<p>Those are different scientific operations and should leave different records.</p>

<h2 id="5-writing-the-design-down-makes-it-reproducible">5. “Writing the design down makes it reproducible”</h2>

<p>Preregistration sounds excessive for an engineering project. I found it useful
for a blunt practical reason: it stopped me from making dozens of individually
reasonable decisions <em>after</em> seeing the data.</p>

<p>For one experiment I froze, in advance: task-selection rules, run count, model
and serving configuration, state representation, divergence and reconvergence
definitions, treatment arms, outcome metric, stop conditions, and the handling
of infrastructure failures.</p>

<p>It did not prevent mistakes. One donor-selection rule later produced two
experimental arms with identical terminal targets, which made the planned
treatment contrast unidentifiable.</p>

<p>What mattered was what happened next. There were other runs in the same dataset
that would have produced different terminal targets, and picking one of those
would have been easy. The gate returned <code class="language-plaintext highlighter-rouge">UNIDENTIFIABLE</code>, Phase B did not run,
the rule was amended and documented, and the amendment was applied only to a new
sample.</p>

<p>Preregistration did not make the original design correct. It made the design
failure visible, which turned out to be the more useful property.</p>

<h2 id="6-repetition-does-not-rescue-a-bad-metric">6. Repetition does not rescue a bad metric</h2>

<p>Suppose I run an agent 100 times and precisely estimate how often it produces
different patches. If all of those patches are correct, I have precisely
measured a behavioral difference that may not matter.</p>

<p>This was the most important conceptual correction in the AgentSeism work.
Initially I focused on trajectory variation: do executions diverge, do they
reconverge, do they finish in different source states? Those are useful
behavioral measurements. They are not task outcomes.</p>

<p>For coding agents the stronger outcome is externally measurable —
<code class="language-plaintext highlighter-rouge">FAIL_TO_PASS</code> and <code class="language-plaintext highlighter-rouge">PASS_TO_PASS</code> give a correctness label independent of the
agent’s own reasoning. That creates two layers: <strong>behavioral variation</strong> and
<strong>consequential variation</strong>. Different trajectories with the same correctness may
be harmless diversity; similar trajectories that repeatedly fail may be stable
unreliability.</p>

<div class="language-text highlighter-rouge"><div class="highlight"><pre class="highlight"><code>             REPEATED EXECUTIONS
                     │
                     ▼
         How much does behavior vary?
                     │
              trajectory metrics
                     ▼
        ┌─────────────────────────┐
        │ Behavioral variation    │
        └────────────┬────────────┘
                     │  + external labels
                     ▼
        ┌─────────────────────────┐
        │ Consequential variation │
        │ Did the variation       │
        │ change correctness?     │
        └─────────────────────────┘
</code></pre></div></div>

<p><strong>Figure 5.</strong> More repetitions reduce uncertainty about the metric you chose.
They do not make the metric more meaningful.</p>

<p>This is the direction of the next experiments: attach correctness labels to
repeated trajectories, and find out which execution differences actually
correlate with — or cause — failure. It is the same question
<a href="/agents-diverge-at-temperature-zero/">Experiment 01</a> ended on.</p>

<h2 id="7-evaluation-measures-the-model">7. “Evaluation measures the model”</h2>

<p>An agent evaluation result belongs to an entire protocol. “Agent X achieves 72%
accuracy” is incomplete without knowing what execution process produced the
number: model and exact revision, sampling configuration, tool definitions,
agent implementation, context-window limit, step limit, retry policy, container
version, serving engine, concurrency policy, termination semantics, scoring
implementation.</p>

<p>Some of those look like infrastructure trivia. They still change the
distribution being measured.</p>

<p>I considered running several AgentSeism trajectories concurrently to cut GPU
rental time. From a serving perspective that is obviously right — continuous
batching improves utilization. But the experiment was studying run-to-run
variation under temperature-zero inference, and changing concurrency changes
batch composition and potentially the numerical execution path. A serving
optimization would have altered the stochastic process under study.</p>

<p>In production I would enable concurrency. For that experiment I kept runs
sequential. The correct configuration depends on the question.</p>

<h2 id="8-the-protocol-i-use-now">8. The protocol I use now</h2>

<p>The pipeline I want looks less like <code class="language-plaintext highlighter-rouge">run → score</code>:</p>

<figure class="wide">
  <img src="/assets/img/evaluation-pipeline.png" alt="Freeze protocol, repeated executions, classify termination into completed, agent failure, infrastructure failure and censored; completed and agent failures flow into validated measurements, external labels, variance estimation and comparison, while infrastructure failures and censored runs are reported separately." loading="lazy" />
  <figcaption><b>Figure 6.</b> Infrastructure failures and censored runs leave the main path, but they are reported rather than dropped.</figcaption>
</figure>

<p>In practice, six questions before I trust an agent-evaluation result:</p>

<ol>
  <li><strong>Is the protocol frozen?</strong> Can I state exactly what system produced these
executions?</li>
  <li><strong>Do I have repeated runs?</strong> Have I measured run-to-run variability rather
than assuming it away?</li>
  <li><strong>Do termination classes have explicit semantics?</strong> Can I separate agent
failure from infrastructure failure and censoring?</li>
  <li><strong>Has the measurement been validated?</strong> Do the hashes, state representations
or judge scores correspond to the construct I claim to measure?</li>
  <li><strong>Is there an external outcome?</strong> Am I measuring task success, or only
behavioral difference?</li>
  <li><strong>Is uncertainty visible?</strong> Does the reported result reflect the distribution
I actually observed?</li>
</ol>

<p>If I cannot answer these, adding more benchmark tasks is usually not the first
thing I need.</p>

<h2 id="9-what-i-still-dont-know">9. What I still don’t know</h2>

<p>This protocol solves part of the problem. Repeated executions get expensive, and
the right number of repetitions depends on the agent’s variance and on the
decision being made. Some evaluations have deterministic external labels; others
need human or model judges, which introduce another layer of measurement noise.
Production agents also operate on non-stationary task distributions, which makes
any fixed benchmark an incomplete proxy for deployed behavior.</p>

<p>The question I care about next is more specific: once repeated executions carry
external correctness labels, can we identify <strong>where successful and failed
trajectories begin to become distinguishable, before termination?</strong></p>

<p>If so, evaluation stops being only retrospective. Instead of “which agent scored
higher?”, the question becomes “when did this execution become more likely to
fail, and could an intervention at that point have changed the outcome?” That is
the bridge between agent evaluation and agent reliability.</p>

<h2 id="conclusion">Conclusion</h2>

<p>The biggest change is that I no longer treat the benchmark score as the
beginning of the analysis. It is the end of a measurement pipeline.</p>

<p>Before trusting the score I want to know what protocol generated it, how much
executions vary, why each run terminated, whether the instrumentation measures
what I think it measures, and whether behavioral differences are grounded in an
external outcome.</p>

<p>So the principle is not simply <em>run the agent more than once</em>. It is:</p>

<blockquote>
  <p><strong>Repeated runs reveal variance. Valid measurement tells you what varied.
External outcomes tell you whether the variation mattered.</strong></p>
</blockquote>

<p>Only then does it seem fair to call the result an evaluation.</p>]]></content><author><name>Ruxi Zhang</name></author><summary type="html"><![CDATA[What repeated agent experiments taught me about variance, censoring, measurement bugs, and knowing whether a result is real — five evaluation assumptions that broke, and the protocol that replaced them.]]></summary><media:thumbnail xmlns:media="http://search.yahoo.com/mrss/" url="https://lyr-ai.github.io/assets/img/evaluation-hero.jpg" /><media:content medium="image" url="https://lyr-ai.github.io/assets/img/evaluation-hero.jpg" xmlns:media="http://search.yahoo.com/mrss/" /></entry></feed>