← Back to writing

Switching to a new LLM? Run a blinded side-by-side eval first

Before you switch a production AI feature to a newly released model, run a blinded side-by-side comparison: send the same real inputs to your current model and the candidate, hide which output came from which, and have qualified reviewers pick the better answer against a written rubric. Benchmarks tell you a model is better in general. A blinded comparison on your own traffic tells you whether it is better for your users, and where it is worse.

This matters right now because model releases are arriving faster than most teams can evaluate them. In early October 2026 Anthropic released Claude Haiku 5.5, Sonnet 5.5, and Opus 5.5, and says Sonnet 5.5 runs about 30% faster and costs up to 30% less for most work than Sonnet 5. Those are real incentives to switch. They are also exactly the kind of headline that leads teams to change a model string on Friday and read support tickets on Monday.

Why benchmark wins do not settle the switch

Vendor and public benchmarks measure broad capability on fixed task sets. Your feature has a narrower job with its own prompt, tools, retrieval context, tone rules, and output format. A new model can score higher on every public benchmark and still:

  • Ignore a formatting instruction your parser depends on.
  • Become more verbose, which raises token cost and breaks UI length limits.
  • Refuse a category of request your old model handled, or accept one it correctly refused.
  • Handle your long-tail inputs differently, such as mixed-language tickets or messy CSV uploads.

Automated regression suites catch some of this. If you already have one, run it first; our guide to catching regressions in CI when models update covers that layer. But regression tests check known failure cases. A side-by-side review answers a different question: across a representative sample of real work, which model produces the output your users would prefer?

How to run a blinded side-by-side comparison

Step 1: sample real inputs, not a demo set

Pull 100 to 300 recent production inputs, with personal data removed or masked. Stratify the sample so it covers your main request types plus the hard tail: long inputs, ambiguous requests, inputs that previously caused complaints. A sample of only clean, typical cases will make every model look the same.

If you do not have production traffic yet, build the set from real support tickets, sales calls, and internal dogfooding. The process in building an eval set from production failures applies directly.

Step 2: generate both outputs under production conditions

Run each input through the full pipeline twice: once with the current model, once with the candidate. Keep everything else identical, including system prompt, retrieval results, tool definitions, and temperature. If the new model needs prompt changes to behave, that is a finding in itself; record it and test the adjusted prompt as a separate candidate rather than mixing the two.

Log latency, input tokens, and output tokens for every call. You will need them for the decision later.

Step 3: blind and shuffle

Strip anything that identifies the model: names, signature phrases in system text, response metadata. Randomize which output appears on the left and which on the right for each pair. Reviewers who know which output is "the new model" tend to favor it, and position bias (preferring the first answer shown) is well documented for both human raters and LLM judges.

Step 4: judge against a rubric, with a reason

Give reviewers the original input, both outputs, and a short rubric drawn from your feature's acceptance criteria: correct, complete, follows format, appropriate tone, safe. Ask for one of four verdicts: A better, B better, tie, or both unacceptable. Require a one-sentence reason for every non-tie.

The "both unacceptable" option is important. It surfaces inputs where neither model works, which are often more valuable to fix than the model choice itself. If you have not written acceptance criteria for the feature yet, start with acceptance criteria and evals for an AI feature.

Step 5: use reviewers who know the domain

The reviewer pool decides how much the result is worth. For a coding assistant, reviewers need to read the language well enough to spot a subtle bug. For a finance or support feature, they need to know what a correct answer looks like in that domain. Have at least two reviewers judge a shared subset of 30 to 50 pairs so you can measure how often they agree; if agreement is low, tighten the rubric before reading anything into the totals.

Reading the results

Three numbers matter, in this order.

Win rate, excluding ties

Of the pairs where reviewers preferred one output, how often did the candidate win? With 200 pairs and, say, 150 non-tie verdicts, a 55% candidate win rate is weak evidence; a 65% or higher win rate is a clearer signal. Report a confidence interval rather than a single percentage, and do not switch on a near coin flip unless cost or latency gains are large.

Losses by category

An overall win can hide a cluster of losses. Break results down by the request types you stratified on. A candidate that wins on simple requests but loses on long documents may still be the right choice, but only if you route long documents to the old model or fix the prompt.

Cost and latency per accepted answer

Combine quality with the token and latency logs from step 2. A model that is 30% cheaper per token but writes 40% longer answers is not cheaper for your feature. Measure cost per accepted answer, not cost per million tokens.

Common ways the comparison goes wrong

  • Using only an LLM judge with no human check. An LLM judge is useful for volume, but it can prefer longer answers or answers in its own style. Calibrate it against human verdicts on the same pairs first; see using an LLM judge as a release gate for how.
  • Comparing a tuned prompt on the old model against an untuned prompt on the new one. Either keep the prompt fixed or treat each prompt and model pair as its own candidate.
  • Reviewing too few pairs. Twenty pairs will not show a 10-point difference in win rate with any confidence.
  • Skipping the rollout plan. A good result justifies a staged rollout with monitoring, not an instant full cutover. If the old model is being retired on a deadline, our model deprecation migration plan covers sequencing.

A realistic timeline

For a single feature, a team that already has logging and a rubric can usually finish this in a few working days: about a day to sample and generate outputs, one to two days of review, and half a day to analyze and decide. The review step is the bottleneck, because it needs people who understand both the domain and what a good answer looks like.

That review step is the work Boundev now focuses on. Our blinded response comparison and code evaluation services give AI teams qualified expert reviewers, rubric scores, and written rationales, so a model switch is a decision backed by evidence rather than a benchmark headline.

Frequently asked questions

How many examples do I need for a side-by-side model comparison?

Plan for 100 to 300 pairs for a single feature. Fewer than about 100 non-tie verdicts makes it hard to tell a real difference from noise unless the gap is large.

Should reviewers know which model produced which output?

No. Hide model identity and randomize left and right position for every pair. Both brand and position bias shift verdicts measurably.

Can I use an LLM judge instead of human reviewers?

You can use one to scale, after you calibrate it against human verdicts on the same pairs. Keep human review for safety-sensitive categories and for a regular audit sample.

What if the new model wins on quality but costs more?

Compute cost per accepted answer and the value of the quality gain for that feature. Many teams route only the hardest request types to the stronger model and keep the rest on the cheaper one.

Work with Boundev

Put expert judgment to work.

Talk to our team about evaluating AI-generated code, comparing responses, and building clear rubrics.