Inter-rater agreement for LLM evals: is your human review reliable?
If two people grade the same 200 AI answers and agree on 85% of them, that sounds like a reliable eval. It often is not. When most answers pass, two reviewers who barely read the output will still agree most of the time by chance. Inter-rater agreement statistics such as Cohen's kappa and Krippendorff's alpha correct for that, and they tell you whether your human labels are solid enough to gate a release, train a judge model, or compare two LLMs.
This post walks through how to measure reviewer agreement on LLM output, which number to use, what thresholds mean, and what to do when agreement comes back low. It is written for teams who already run human review on AI features and want to know whether they can trust it.
Why raw percent agreement misleads
Percent agreement is simply the share of items where reviewers gave the same label. It ignores how often they would agree if they were guessing in line with their own habits. On LLM evals that gap is large, because most production answers are acceptable and the failures you care about are rare.
A worked example
Two reviewers grade 200 support-bot answers as pass or fail. The counts are:
- Both mark pass on 150 answers.
- Both mark fail on 20 answers.
- Reviewer A passes and reviewer B fails on 18 answers.
- Reviewer A fails and reviewer B passes on 12 answers.
Raw agreement is (150 + 20) / 200 = 0.85. Now compute the agreement you would expect by chance. Reviewer A passes 168 of 200 answers (0.84) and reviewer B passes 162 of 200 (0.81). Chance agreement is 0.84 x 0.81 plus 0.16 x 0.19, which is 0.6804 + 0.0304 = 0.711.
Cohen's kappa is (observed - chance) / (1 - chance) = (0.85 - 0.711) / (1 - 0.711) = 0.48. The 85% headline shrinks to a kappa of 0.48, which is moderate at best. And the disagreement sits exactly where it hurts: of the 50 answers either reviewer failed, they agreed on only 20. If those failures are hallucinated refund policies, your eval is not catching them consistently.
The prevalence trap
Kappa drops when one label dominates, even when reviewers look careful. This is sometimes called the kappa paradox. It is not a bug in the statistic. It is telling you that on a skewed set, agreement on the common label is cheap and agreement on the rare label is what matters. Report percent agreement and kappa together, and look at agreement on the rare class directly.
Which agreement statistic to use
Pick the statistic from the shape of your review process, not from what the last paper used.
Cohen's kappa: exactly two reviewers
Use Cohen's kappa when the same two people label every item with a categorical label such as pass/fail or a short failure taxonomy. It is easy to compute and easy to explain. For ordered scales, such as a 1 to 5 helpfulness score, use weighted kappa so that a 4 versus 5 disagreement costs less than a 1 versus 5.
Fleiss' kappa: a fixed number of reviewers per item
Fleiss' kappa extends the idea to three or more raters, where each item gets the same number of ratings but not necessarily from the same people. It suits a review pool where any three of eight reviewers grade each answer.
Krippendorff's alpha: the general case
Krippendorff's alpha handles any number of raters, missing ratings, and nominal, ordinal, interval or ratio data through a choice of distance function. Real eval programs have missing ratings all the time: a reviewer skips an item, the pool changes mid-week, a second pass covers only a subset. If you are unsure, alpha is the safe default. Always report which distance metric you used, because nominal and ordinal alpha on the same data can differ a lot.
What thresholds actually mean
The most quoted convention comes from Krippendorff: alpha of 0.800 or above is treated as reliable, and values from 0.667 to 0.800 support only tentative conclusions. Krippendorff himself described 0.667 as a lowest limit, not a target. For kappa, many teams still cite the Landis and Koch (1977) bands, where 0.41 to 0.60 is moderate and 0.61 to 0.80 is substantial.
Treat these as conventions, not laws. Two practical adjustments matter more:
- Set the bar by the decision. A release gate that can block a launch, or labels that will train a judge model, need a higher bar than an exploratory look at a new prompt.
- Compare against the task's ceiling. Some judgments are hard even for experts. Published NLP annotation work regularly reports human-to-human agreement well under 0.8 on subjective tasks, so an LLM judge that matches humans as well as humans match each other may be the best you can get on that task.
That second point is how the original LLM-as-a-judge study framed its result: Zheng et al. (2023) reported that GPT-4 agreed with human preferences over 80% of the time, about the same rate at which humans agreed with each other. Human-human agreement is the yardstick for any automated grader. We cover the gate side of this in our post on using an LLM judge as a release gate.
A practical agreement workflow for LLM evals
Here is the process we use when a team asks whether its human review is trustworthy.
1. Double-label a calibration slice
You do not need every item labeled twice. Pick a stratified slice of 100 to 200 items that covers each feature area and over-samples the failure types you care about. Have at least two reviewers label it independently, without seeing each other's labels or which model produced the output.
2. Write the rubric before you score
Low agreement is usually a rubric problem, not a reviewer problem. Vague criteria such as "helpful" or "accurate enough" produce noise. Write pass and fail anchors with real examples for each criterion. Our rubric for grading AI-generated code shows the level of specificity that works.
3. Compute agreement per criterion, not just overall
One overall alpha hides the criterion that is broken. Teams often find factual accuracy at 0.85 and tone at 0.40 in the same rubric. The fix for tone is a better definition or dropping it from the gate, not more reviewers.
4. Adjudicate and read the disagreements
For every disagreement, have a third reviewer or the rubric owner decide, and write one line on why. After 30 or 40 of these, patterns appear: an ambiguous rule, a missing edge case, a reviewer who reads the policy differently. Update the rubric and re-score a fresh slice.
5. Recalibrate on a schedule
Agreement drifts as reviewers change, products change, and the model's failure modes change. Re-run the calibration slice every model change and at least monthly. Add new failures from production to the slice, as described in building an eval set from production failures.
Using agreement to validate an LLM judge
Once human labels are reliable, the same math validates an automated grader. Add the judge model as one more rater and compute alpha across the humans plus the judge, or compute judge-to-adjudicated-label agreement and compare it with human-to-human agreement on the same slice.
Watch three things:
- Agreement on the rare class. A judge that passes everything will look fine on a mostly-passing set. Check how many human-flagged failures it also flags.
- Known judge biases. Position bias, length bias and a preference for its own model family were all documented in the Zheng et al. paper. Shuffle order and test pairs where the shorter answer is correct.
- Stability across model versions. A judge that agreed with humans last quarter may not after a provider update, so keep the judge in your eval regression suite.
The same discipline applies to blinded model comparisons. If reviewers cannot agree on which of two answers is better, a 54% win rate for the new model is noise. Our guide to blinded side-by-side evals before switching models explains how to report win rates with that in mind.
What to do when agreement is low
- Below about 0.4: stop using these labels for decisions. Rewrite the rubric, retrain reviewers on anchored examples, and re-score.
- Roughly 0.4 to 0.67: usable for direction, not for gates. Adjudicate every disagreement and look for the one or two criteria dragging the score down.
- 0.67 to 0.8: fine for most product decisions. For a launch gate, require adjudication on items near the threshold.
- Above 0.8: solid. Keep recalibrating so it stays there.
If the task is hard by nature, such as judging medical or legal nuance, you may need domain experts rather than general reviewers. That is the gap Boundev's expert AI evaluation network is built to fill.
FAQ
How many items do I need to measure inter-rater agreement?
For a working estimate, 100 to 200 double-labeled items is usually enough to spot a broken rubric. If the failure class you care about is rare, over-sample it so you have at least a few dozen examples of it, or the agreement number will say little about failures.
Is Cohen's kappa or Krippendorff's alpha better for LLM evaluation?
Use Cohen's kappa when exactly two reviewers label every item. Use Krippendorff's alpha when you have more reviewers, missing ratings, or ordinal scales. Most real eval programs fit alpha better.
What is a good kappa score for human review of AI output?
Conventions treat 0.61 to 0.80 as substantial agreement and above 0.80 as strong. For a release gate, aim for at least 0.67 on alpha and preferably 0.8, and always compare with what human reviewers achieve on the same task.
Can an LLM judge replace human reviewers?
Only for criteria where it agrees with adjudicated human labels about as well as humans agree with each other, including on the rare failure class. Keep a human-labeled calibration slice and re-check the judge whenever the judge or the graded model changes.
Why is my kappa low when reviewers agree 90% of the time?
Because most of your items share one label, so chance agreement is already high. Kappa measures agreement beyond chance. Check agreement on the minority label, which is usually where the real disagreement sits.
Put expert judgment to work.
Talk to our team about evaluating AI-generated code, comparing responses, and building clear rubrics.