Define the decision first
A release comparison, a review-comment audit, and a training-data project need different judgments. Begin with the decision the buyer will make. Add a criterion only when its result could affect that decision or explain a meaningful difference.
For a code patch, four starting dimensions are correctness, instruction following, maintainability, and risk. These are proposed design choices, not a universal standard. Adapt them to the actual task.
Give each criterion a boundary
| Dimension | Review question | Evidence to request |
|---|---|---|
| Correctness | Does the change implement the required behavior? | Contract, diff, context, and relevant client-supplied execution results |
| Instruction following | Were scope and constraints respected? | Original prompt and explicit constraints |
| Maintainability | Can the change be understood and maintained in context? | Surrounding code and project conventions |
| Risk | Does the supplied evidence show a material concern? | Affected behavior, assumptions, and specialist evidence when needed |
Write anchors people can apply
For a five-point scale, define the ends and the middle in observable terms. For example: 1 means the central requirement is contradicted; 3 means the main path is addressed but a material requirement is incomplete; 5 means the supplied evidence supports the full stated contract. Add examples for scores 2 and 4 before calibration.
These anchors do not turn a 5 into a guarantee of production reliability. Use a separate “insufficient evidence” result rather than treating missing evidence as zero. Capture the reason and the evidence needed to continue.
Keep blockers visible
Define critical failures separately from averages. If the agreed task requires rejecting revoked sessions, accepting one should not be hidden by strong readability scores. Report the blocker and the criterion it violates.
For pairwise comparisons, allow A preferred, B preferred, equivalent, and insufficient evidence. A forced winner can turn a tie or missing context into a misleading preference label.
Calibrate before a large batch
Give reviewers the same small set and collect their independent reasons. Discuss disagreements after submission, revise ambiguous anchors, and version the rubric. Use a fresh set to check whether the revised instructions are understood. Record real calibration outcomes before claiming reviewer consistency.
