EVALUATION DESIGN

Build a sample that matches the decision.

Make the population, selection method, and exclusions explicit before interpreting an evaluation result.

Name the population

Describe the work your result is meant to represent: a particular repository, language, task category, customer workflow, or model release. A sample assembled from convenient examples may be useful for debugging, but its conclusions should remain limited to those examples.

SWE-bench distinguishes multiple benchmark sets with different task scopes. That distinction is a useful reminder to identify the dataset and evaluation setting when comparing results.

Separate discovery from comparison

For finding failure modes, deliberately difficult examples can be useful. For estimating how frequently a problem occurs, you need a sampling method that supports that estimate. Do not report a challenge-set failure rate as an overall production rate.

For a model comparison, use the same eligible tasks and comparable supplied context. Record variant identifiers privately, randomize presentation where appropriate, and keep a mapping outside the review surface. Reviewers should judge the work against the rubric.

Choose the coverage you need

Keep the sampling frame and selection logic with the dataset. A count of reviewed items alone does not explain whether the sample covers the intended work.

  • Common workflows that account for meaningful use.
  • Important boundary and error-handling cases.
  • Task sizes and languages relevant to the buyer.
  • Cases where current automated checks provide incomplete evidence.
  • Explicit exclusions, including material that cannot be shared lawfully or safely.

Version what could change the answer

Record the prompt, artifact revision, rubric version, and any supplied execution environment. If requirements change during review, document which items are affected. Avoid quietly mixing the old and new criteria in one summary.

State the number of eligible items, completed reviews, missing reviews, and unresolved adjudications. Report exclusions and their reasons. When generalizing beyond the sample, use a method and uncertainty estimate appropriate to the sampling design; a small convenience sample may support no such estimate.

A useful first brief

Write one sentence describing the decision, one describing the population, and one describing the selection method. Add a short list of exclusions and what evidence each reviewer receives. This is often a better starting point for a scoped pilot than asking for an arbitrary large number of labels.

Further reading

YOUR NEXT EVALUATION

Start with the decision.

Share a short brief. Scope, availability, and fees are agreed before work starts.

Talk to Mayur ↗