When to replace regex checks with an LLM judge (and how to test the judge)
Most LLM features start with checks like this:
const REFUSAL = /\b(I can't|I cannot|I'm unable to|as an AI)\b/i
if (REFUSAL.test(reply)) flag(reply, 'refused')That works for a week. Then the model starts saying "That's outside what I can help with," and you add another pattern. Then a perfectly good answer that says "I can't stress this enough" gets flagged. A few months later the file has dozens of patterns and nobody trusts what it reports.
The regex is fine, as regexes go. The trouble is the question it's trying to answer. "Did the model refuse?" is about meaning, and a regex only sees characters.
An LLM judge is a second model call that answers that question directly. It can work much better. It also costs money, adds latency, and can be wrong in ways that are harder to see. This post covers when the switch is worth it and how to test the judge before you rely on it.
Where regex and code checks win
Keep deterministic checks for anything that is about form, not meaning:
- Format. Is it valid JSON? Does it match the schema? Is it under the length limit?
- Exact strings. Does the output contain an API key pattern, an internal hostname, or a banned word from a fixed list?
- Structure. Does every citation marker like
[3]point at a source that exists? - Hot paths. Anything that runs on every request where an extra model call would hurt latency.
These checks are free, they run in microseconds, they never change behavior on their own, and when one fires you can read the pattern and see exactly why. An LLM judge gives up all four of those properties. Don't trade them away for a check that a schema validator does perfectly.
Where a judge wins
Reach for a judge when the check is about meaning and you can see the regex failing:
- Did the reply refuse, deflect, or change the subject?
- Does the answer actually use the retrieved sources, or just mention them?
- Is the tone appropriate for a customer-facing message?
- Did the summary leave out the one fact that mattered?
A good sign that you need a judge: you keep adding patterns, and each one fixes one example and breaks another.
The setup I use is layered. Deterministic gates run first and fail fast. Only outputs that pass them reach the judge. The judge's verdicts get sampled into a review queue, and those reviews become new labeled examples.
Writing the judge
Keep each judge narrow: one criterion per judge, a yes/no verdict, and a short reason. A judge asked to "rate quality from 1 to 10" gives you numbers that drift and are hard to act on. A judge asked "does this reply refuse the user's request?" gives you something you can measure.
I use a forced tool call so the verdict always comes back in a fixed shape:
import Anthropic from '@anthropic-ai/sdk'
const client = new Anthropic()
const JUDGE_MODEL = process.env.JUDGE_MODEL ?? ''
export type Verdict = { reason: string; fails: boolean }
const verdictTool: Anthropic.Tool = {
name: 'verdict',
description: 'Record your verdict on the output.',
input_schema: {
type: 'object',
properties: {
reason: { type: 'string', description: 'One sentence, written before deciding.' },
fails: { type: 'boolean' },
},
required: ['reason', 'fails'],
},
}
export async function judge(criterion: string, output: string): Promise<Verdict> {
const res = await client.messages.create({
model: JUDGE_MODEL,
max_tokens: 300,
system: `You check exactly one thing: ${criterion}\nAnswer only about that.`,
tools: [verdictTool],
tool_choice: { type: 'tool', name: 'verdict' },
messages: [{ role: 'user', content: `<output>\n${output}\n</output>` }],
})
const call = res.content.find(b => b.type === 'tool_use')
if (!call) throw new Error('judge returned no verdict')
return call.input as Verdict
}Putting reason before fails in the schema nudges the model to explain first and decide second. Setting tool_choice to a specific tool means you never have to parse free text. Wrapping the output in tags makes it harder for the judged text to pass as instructions to the judge.
Pin the judge model to a specific version, not an alias that moves. If the judge changes under you, every number you measured before is stale.
Test the judge like any other classifier
A judge is a classifier, so you test it like one: against examples a human has already labeled.
Build the labeled set first. Pull real outputs from logs and have a person mark each one pass or fail for the criterion. A hundred or two hundred examples is enough to start. Make sure it includes the hard cases: the polite refusal, the answer that mentions "can't" in passing, the reply that is half helpful. A set made only of easy examples will tell you the judge is great.
Run both checks on the same set. Run the regex and the judge on every labeled example and compare each against the human label. Now the two are on equal footing.
Look at more than accuracy. If only a small share of outputs fail, a check that always says "pass" scores high on accuracy and catches nothing. Look at these instead:
- Precision: when the check flags something, how often is it right? Low precision means alert fatigue.
- Recall: of the real failures, how many did it catch? Low recall means a false sense of safety.
- Cohen's kappa: agreement with the human, adjusted for the agreement you would get by luck. It's useful when one class is rare.
The scoring code is short:
// "Positive" means the check flagged a failure.
export type Pair = { human: boolean; check: boolean }
export function score(pairs: Pair[]) {
let tp = 0, fp = 0, fn = 0, tn = 0
for (const { human, check } of pairs) {
if (check && human) tp++
else if (check && !human) fp++
else if (!check && human) fn++
else tn++
}
const n = pairs.length
const agreement = (tp + tn) / n
// Agreement you would get by chance, given how often each side says "fail".
const chance = ((tp + fp) * (tp + fn) + (fn + tn) * (fp + tn)) / (n * n)
return {
agreement,
precision: tp / (tp + fp),
recall: tp / (tp + fn),
kappa: (agreement - chance) / (1 - chance),
confusion: { tp, fp, fn, tn },
}
}With 100 examples where the check gets 8 true positives, 2 false alarms, 2 misses and 88 correct passes, this returns 0.96 agreement, 0.8 precision, 0.8 recall and a kappa of about 0.78. The 0.96 looks great on a slide. Kappa is the number I'd put on the slide instead.
Read the disagreements. The metrics tell you how often the judge is wrong. The disagreements tell you why. Sometimes the judge is wrong. Sometimes the human label is wrong, and you want to know that too. Sometimes the criterion is vague and two people would disagree as well. Fix the wording and label again.
Known judge biases
The MT-Bench paper (Zheng et al., 2023) documents biases that LLM judges show: a preference for whichever answer comes first in pairwise comparisons, a preference for longer answers, and a tendency to favor text written by the same model. Binary, single-output judges avoid the position problem. Length bias still shows up, so put long and short examples in your labeled set and check the judge on both.
Latency and cost
A judge adds a model call per checked output. Before you put it inline, decide where it runs:
| Where | When it fits |
|---|---|
| Offline eval only | Prompt changes, model upgrades, regression suites |
| Sampled in prod | Watching quality trends on a share of real traffic |
| Inline, blocking | Only when a bad output is worse than a slow response |
Most judges belong in the first two rows. A blocking judge on every request doubles your model calls and adds a second source of errors to the request path. Use a smaller, cheaper model for the judge when your labeled set shows it scores about the same as a bigger one. That's one more thing the labeled set lets you measure instead of guess.
Takeaways
- Keep deterministic checks for format, exact strings and anything on the hot path. They're free and you can read them.
- Switch to a judge when the question is about meaning and the regex keeps needing new patterns.
- One criterion per judge, a forced structured verdict, a reason before the verdict, and a pinned model version.
- Test the judge against a human-labeled set with hard cases. Report precision, recall and kappa, not just accuracy.
- Re-run the labeled set whenever the judge prompt or model changes.
Anthropic's guide on building evals and OpenAI's evals guide go further into model-graded checks if you want more depth.