RUBRIC DESIGN

Choose review dimensions that answer your question.

A compact rubric should make a decision easier, with explicit anchors and a way to handle missing evidence.

Define the decision first

A release comparison, a review-comment audit, and a training-data project need different judgments. Begin with the decision the buyer will make. Add a criterion only when its result could affect that decision or explain a meaningful difference.

For a code patch, four starting dimensions are correctness, instruction following, maintainability, and risk. These are proposed design choices, not a universal standard. Adapt them to the actual task.

Give each criterion a boundary

DimensionReview questionEvidence to request
CorrectnessDoes the change implement the required behavior?Contract, diff, context, and relevant client-supplied execution results
Instruction followingWere scope and constraints respected?Original prompt and explicit constraints
MaintainabilityCan the change be understood and maintained in context?Surrounding code and project conventions
RiskDoes the supplied evidence show a material concern?Affected behavior, assumptions, and specialist evidence when needed

Write anchors people can apply

For a five-point scale, define the ends and the middle in observable terms. For example: 1 means the central requirement is contradicted; 3 means the main path is addressed but a material requirement is incomplete; 5 means the supplied evidence supports the full stated contract. Add examples for scores 2 and 4 before calibration.

These anchors do not turn a 5 into a guarantee of production reliability. Use a separate “insufficient evidence” result rather than treating missing evidence as zero. Capture the reason and the evidence needed to continue.

Keep blockers visible

Define critical failures separately from averages. If the agreed task requires rejecting revoked sessions, accepting one should not be hidden by strong readability scores. Report the blocker and the criterion it violates.

For pairwise comparisons, allow A preferred, B preferred, equivalent, and insufficient evidence. A forced winner can turn a tie or missing context into a misleading preference label.

Calibrate before a large batch

Give reviewers the same small set and collect their independent reasons. Discuss disagreements after submission, revise ambiguous anchors, and version the rubric. Use a fresh set to check whether the revised instructions are understood. Record real calibration outcomes before claiming reviewer consistency.

YOUR NEXT EVALUATION

Start with the decision.

Share a short brief. Scope, availability, and fees are agreed before work starts.

Talk to Mayur ↗