Skip to content
Awsaf Alam
GitHubLinkedIn

Record and replay: deterministic evals for LLM agents

— AI Agents, Testing, Evals — 7 min read

The first test suite for an agent usually calls the real model. It works, for a while. Then it takes four minutes, costs real money on every push, and fails one run in five for reasons nobody can reproduce. People start re-running CI until it goes green, and at that point the suite has stopped telling you anything.

The fix is old. HTTP testing libraries like Ruby's VCR, Python's vcrpy and Netflix's Polly.js have done it for years: record real responses once, save them to a file, and replay the file in tests. For agents, you record two kinds of calls: model calls and tool calls.

This post walks through a small cassette implementation in TypeScript, how to write golden cases on top of it, and how to catch the drift that replay hides.

What a cassette is

A cassette is a file that maps a request to the response it got. In record mode, calls go to the real model and real tools, and every request/response pair is written down. In replay mode, nothing touches the network. Each request is looked up on the tape, and if it isn't there, the test fails.

Two lanes. In record mode, started by hand with RECORD=1, the agent calls through a cassette wrapper to the live model and live tools, and the cassette writes case.json. In replay mode, on every CI run, the agent calls the same wrapper, which reads responses from case.json. A request not on the tape fails the test without using the network.

That last part matters most. A replay miss should never quietly fall through to the live API. If it does, your "offline" suite is online again whenever a prompt changes, and you're back to slow, flaky, paid tests without noticing.

A small implementation

This wraps any async function. It has no dependencies beyond Node's standard library.

ts
import { createHash } from 'node:crypto'
import { existsSync, mkdirSync, readFileSync, writeFileSync } from 'node:fs'
import { dirname } from 'node:path'
 
type Mode = 'record' | 'replay'
 
export class Cassette {
  private entries: Record<string, unknown> = {}
  private counts = new Map<string, number>()
 
  constructor(private path: string, private mode: Mode) {
    // Record mode starts empty, so a re-record never keeps stale entries.
    if (mode === 'replay') {
      if (!existsSync(path)) throw new Error(`No cassette at ${path}. Record it with RECORD=1.`)
      this.entries = JSON.parse(readFileSync(path, 'utf8'))
    }
  }
 
  wrap<A, R>(name: string, fn: (args: A) => Promise<R>): (args: A) => Promise<R> {
    return async args => {
      const base = `${name}:${hash(args)}`
      // The same request can legitimately happen twice (polling, retries),
      // so number each occurrence.
      const n = (this.counts.get(base) ?? 0) + 1
      this.counts.set(base, n)
      const key = `${base}#${n}`
 
      if (this.mode === 'replay') {
        if (!(key in this.entries)) {
          throw new Error(`Cassette miss for ${key} in ${this.path}. The request changed; re-record.`)
        }
        return structuredClone(this.entries[key]) as R
      }
      const result = await fn(args)
      this.entries[key] = result
      return result
    }
  }
 
  save(): void {
    if (this.mode !== 'record') return
    mkdirSync(dirname(this.path), { recursive: true })
    writeFileSync(this.path, JSON.stringify(this.entries, null, 2) + '\n')
  }
}
 
// Sort object keys so {a, b} and {b, a} hash the same.
function sortKeys(v: unknown): unknown {
  if (Array.isArray(v)) return v.map(sortKeys)
  if (v && typeof v === 'object') {
    return Object.fromEntries(
      Object.keys(v)
        .sort()
        .map(k => [k, sortKeys((v as Record<string, unknown>)[k])]),
    )
  }
  return v
}
 
function hash(v: unknown): string {
  return createHash('sha256').update(JSON.stringify(sortKeys(v))).digest('hex').slice(0, 16)
}

Three details matter here:

  • Keys come from the request content. The model name, system prompt, messages, tool definitions and parameters all go into the hash. Change any of them and you get a miss, which is what you want: a different request deserves a fresh recording.
  • Repeated requests are numbered. Agents sometimes make the same call twice, for example when polling a job's status. Each occurrence gets its own slot, so the second poll replays the second response.
  • Keep volatile values out of the hash. Request IDs, timestamps or random seeds in the arguments give you a new key every run. Strip them before the call reaches the wrapper, or pass a clock into the agent so "now" is fixed in tests.

Wiring it into the agent

Replay only works if the agent receives its model client and tools as arguments instead of importing them directly. If your agent calls a global client deep inside a helper, fix that first. It's the only refactor this approach needs.

ts
import Anthropic from '@anthropic-ai/sdk'
import { test, expect } from 'bun:test'
import { Cassette } from './cassette'
import { runAgent } from '../src/agent'
import { searchDocs } from '../src/tools'
 
const client = new Anthropic()
const mode = process.env.RECORD ? 'record' : 'replay'
 
test('refund question cites the refund policy', async () => {
  const tape = new Cassette('test/cassettes/refund-question.json', mode)
  const out = await runAgent('Can I get a refund after 30 days?', {
    createMessage: tape.wrap('model', (p: Anthropic.MessageCreateParamsNonStreaming) =>
      client.messages.create(p),
    ),
    searchDocs: tape.wrap('tool.searchDocs', searchDocs),
  })
  tape.save()
 
  expect(out.status).toBe('done')
  expect(out.toolsCalled).toContain('searchDocs')
  expect(out.answer).toMatch(/refund policy/i)
})

runAgent, searchDocs and the shape of out belong to your code. The cassette doesn't care what it wraps.

Record tool calls as well as model calls. If you replay the model but run tools live, a change in the search index changes the tool result, the next model request no longer matches the tape, and you get a miss that has nothing to do with your change.

Golden cases: assert on outcomes, not text

A golden case is a fixed input plus checks on what the agent should do with it. Keep 20 to 50 of them, chosen to cover the paths that matter: the happy path, a tool error, an ambiguous request, a request that should be refused.

Don't snapshot the whole answer. Under replay the text is fixed, so a full snapshot passes. But the next re-record changes the wording, and you end up reviewing a diff of synonyms. Assert on the things that would be bugs:

  • the final status and stop reason
  • which tools were called, and roughly in what order
  • key facts or IDs in the answer
  • things that must never appear (another customer's data, a made-up URL)

For checks about meaning, you can call an LLM judge here too. Record the judge's calls on the same cassette and its verdicts replay just as deterministically.

Replay hides drift, so look for it on purpose

Replay tests answer one question: "did my code change break the agent, given these exact model responses?" They can't tell you whether the model still behaves the way the tape says. Providers update models, aliases move to new versions, and old snapshots get deprecated. Your tape stays the same through all of it.

So run two loops:

Every push to a branch replays the golden cases from cassettes, which takes seconds, costs nothing and needs no network. Separately, a nightly job runs the same golden cases live against the real model and tools. If an outcome changed, that is drift, and the fix is to re-record and review the cassette diff.
  • Every push: replay. Fast, free, the same result every time. This is the gate that blocks merges.
  • Nightly or weekly: run the same golden cases live, in record mode, into a scratch directory. Compare the outcomes with the replay results. If the assertions still pass, the new recordings can replace the old ones. If they fail, a person looks before anything is overwritten.

Pin model versions in config so that "the model changed" is a commit you can see, not something that happened on the provider's side overnight. When you upgrade, re-record everything in the same pull request and review the cassette diff like code.

When re-recording, review the diff

Cassettes are JSON, so they diff well. A re-record after a prompt change should show model responses changing and tool calls mostly staying the same. If a re-record suddenly drops a tool call or adds three new ones, read that part of the pull request first. Don't merge it unread.

Two practical notes. Scrub secrets and personal data before writing a cassette, since it gets committed. And keep cassettes small: one file per golden case is easier to review and re-record than one big shared tape.

Takeaways

  • Wrap both model calls and tool calls. Key them on a hash of the request with its keys sorted.
  • In replay mode, a miss fails the test. It never falls back to the live API.
  • Pass the client and tools into the agent so tests can swap in wrapped versions.
  • Assert on outcomes: status, tools called, key facts. Not the full text.
  • Run golden cases live on a schedule to catch model drift, and review every re-recorded cassette like code.
© 2026 Awsaf Alam