REVIEW QUALITY

Make reviewer disagreement useful.

A disagreement can reveal ambiguous requirements, weak score anchors, or missing evidence. Preserve enough detail to tell which.

Keep the first judgments independent

Ask reviewers to submit their initial conclusion and evidence before showing them another person’s answer. If everyone sees a first opinion at the outset, apparent agreement can reflect influence rather than independent interpretation.

Review tools can support this workflow. LangSmith’s annotation-queue documentation describes reviewer assignment and controls that keep other reviewers’ feedback hidden. A tool setting still needs an agreed operating process.

Classify the cause

CauseUseful next action
Different requirements assumedAsk the requirement owner to clarify the contract
Different score anchorsCalibrate the rubric on an example with a written rationale
Missing contextRequest the specific artifact or execution evidence needed
Evidence interpreted differentlyAsk an adjudicator to compare both reasons against the criterion
Different importance assignedClarify the risk or release-blocking policy

Record more than a final score

Retain the original scores, reasons, artifact version, rubric version, and the final decision. The adjudicator’s note should explain what evidence changed the outcome or why one interpretation was selected. Do not overwrite the original review to make a dataset look unanimous.

Treat “insufficient evidence” as a meaningful outcome. A reviewer should be able to abstain from a conclusion when the supplied materials cannot support it. Forcing a numeric score in every case produces false precision.

Measure with a defined denominator

Report how many items received the required independent reviews, how many were escalated, and how many remain unresolved. State whether agreement is exact category agreement, agreement within a score tolerance, or another defined measure.

A three-item demonstration cannot establish reliable population-level reviewer agreement. Compare like-for-like tasks and rubric versions, and preserve the sampling method. Investigate systematic disagreements before interpreting a single aggregate number.

Use the result to improve the next batch

Add a clarified anchor or evidence requirement to the next rubric version. If a change affects earlier conclusions, record whether affected items will be reviewed again. Keep this change history visible to the buyer so an apparent trend is not mistaken for a model improvement.

Further reading

YOUR NEXT EVALUATION

Start with the decision.

Share a short brief. Scope, availability, and fees are agreed before work starts.

Talk to Mayur ↗