Keep the first judgments independent
Ask reviewers to submit their initial conclusion and evidence before showing them another person’s answer. If everyone sees a first opinion at the outset, apparent agreement can reflect influence rather than independent interpretation.
Review tools can support this workflow. LangSmith’s annotation-queue documentation describes reviewer assignment and controls that keep other reviewers’ feedback hidden. A tool setting still needs an agreed operating process.
Classify the cause
| Cause | Useful next action |
|---|---|
| Different requirements assumed | Ask the requirement owner to clarify the contract |
| Different score anchors | Calibrate the rubric on an example with a written rationale |
| Missing context | Request the specific artifact or execution evidence needed |
| Evidence interpreted differently | Ask an adjudicator to compare both reasons against the criterion |
| Different importance assigned | Clarify the risk or release-blocking policy |
Record more than a final score
Retain the original scores, reasons, artifact version, rubric version, and the final decision. The adjudicator’s note should explain what evidence changed the outcome or why one interpretation was selected. Do not overwrite the original review to make a dataset look unanimous.
Treat “insufficient evidence” as a meaningful outcome. A reviewer should be able to abstain from a conclusion when the supplied materials cannot support it. Forcing a numeric score in every case produces false precision.
Measure with a defined denominator
Report how many items received the required independent reviews, how many were escalated, and how many remain unresolved. State whether agreement is exact category agreement, agreement within a score tolerance, or another defined measure.
A three-item demonstration cannot establish reliable population-level reviewer agreement. Compare like-for-like tasks and rubric versions, and preserve the sampling method. Investigate systematic disagreements before interpreting a single aggregate number.
Use the result to improve the next batch
Add a clarified anchor or evidence requirement to the next rubric version. If a change affects earlier conclusions, record whether affected items will be reviewed again. Keep this change history visible to the buyer so an apparent trend is not mistaken for a model improvement.
