How to write a rubric for grading AI-generated code
A rubric for grading AI-generated code is a short list of criteria, each with a fixed scale and written anchors, that tells a reviewer exactly what a 1, 3, or 5 looks like. Without one, two senior engineers will score the same patch differently and your eval numbers become noise. With one, you can compare models, prompts, and agent versions on the dimensions that actually matter: does the change solve the task, is it safe, and could your team maintain it.
This guide covers the criteria worth scoring, how to write anchors reviewers agree on, how to check that agreement, and the mistakes that quietly ruin code evals.
Why passing tests is not enough
Most teams start grading generated code with one signal: do the tests pass. That is a good floor and a poor ceiling. A patch can pass every test and still be wrong in ways that matter in production:
- It special-cases the exact inputs the tests use instead of fixing the underlying logic.
- It swallows an exception so the failing assertion never runs.
- It adds a dependency, a raw SQL string, or a hardcoded secret that a reviewer would reject on sight.
- It solves the problem but rewrites 400 lines to do it, which no one will approve in a real pull request.
Public benchmarks ran into the same issue. OpenAI released SWE-bench Verified in 2024 after human reviewers found that many original SWE-bench tasks had underspecified issues or unfair tests, and kept a 500-task subset that annotators confirmed was solvable and fairly tested. The lesson for your own evals: tests are evidence, not the verdict. A human or rubric-guided judgment has to sit on top.
The five criteria worth scoring
Keep the rubric small. Every criterion you add slows reviewers down and lowers agreement. Five is a practical maximum for code; three is fine to start.
1. Correctness against the brief
Does the change do what the task asked, including edge cases the task implies but does not list? Score this by reading the brief first, then the diff, then the tests. A patch that is correct only for the tested inputs gets a low score here even if CI is green.
2. Security
Look for injection risks, unsafe deserialization, secrets in code, missing authorization checks, and new dependencies with no clear reason. The OWASP Top 10 is a reasonable checklist for web code. Security is usually best scored as a gate (pass, or fail with a reason) rather than a 1 to 5 scale, because a single real vulnerability should not be averaged away by good style.
3. Maintainability
Would your team accept this diff in review? Consider naming, size of the change relative to the problem, duplication, and whether it follows the patterns already in the codebase. This is the most subjective criterion, so it needs the most concrete anchors.
4. Scope discipline
Did the change stay inside the task? Generated code often edits unrelated files, reformats whole modules, or renames things nobody asked about. Scope creep makes real review harder and hides risk.
5. Test quality
If the model wrote or changed tests, do they check behavior, or were they weakened to pass? A deleted assertion or a test rewritten to match the buggy output should score at the bottom of the scale.
How to write anchors reviewers agree on
A scale of 1 to 5 with no description is a vibe check. Anchors turn it into a measurement. For each criterion, write what a 1, 3, and 5 look like in concrete, observable terms, and let 2 and 4 sit between them.
Here is an example for maintainability:
- Score 5: the diff is the smallest reasonable change, follows existing patterns in the file, and needs no review comments beyond nits.
- Score 3: the change works and is readable, but a reviewer would request one or two structural changes, such as extracting a helper or removing duplication.
- Score 1: a reviewer would reject the diff outright, for example because it rewrites unrelated code, duplicates large blocks, or ignores the conventions used elsewhere in the module.
Good anchors share three traits. They describe things a reviewer can point at in the diff. They avoid words like "good" or "clean" on their own. And they include at least one real example from your own codebase, attached to the rubric as a reference case.
Require a written rationale
Ask reviewers for one or two sentences explaining each score below the top. Rationales do three jobs: they make scores auditable, they let you find where the rubric is unclear, and they become training material for future reviewers or for an LLM judge that you later calibrate against humans.
Check that the rubric actually works
Before you trust rubric scores, measure whether two reviewers agree. A simple process:
- Pick 30 to 50 real patches that cover easy, hard, and borderline cases.
- Have two qualified reviewers score them independently, without seeing each other's scores.
- Compute agreement per criterion. Cohen's kappa is the standard statistic for two raters on categorical scores; a weighted kappa suits ordinal 1 to 5 scales because it treats a 4 vs 5 disagreement as smaller than a 1 vs 5.
- For every disagreement of two points or more, read both rationales and fix the anchor that caused it.
Repeat until agreement is stable. Low agreement almost always means an anchor is vague, not that reviewers are careless. Correctness and security usually reach good agreement quickly; maintainability usually needs a second pass.
If you plan to use an LLM as a judge, this human-scored set is also your calibration set. Score the same patches with the judge and compare. Our guide to using an LLM judge as a release gate covers how to decide when a judge is reliable enough to stand in for people.
Mistakes that quietly ruin code evals
Averaging a security failure away
If security is one of five averaged criteria, a patch with an SQL injection can still score 4 out of 5 overall. Keep hard failures as gates that override the total.
Letting reviewers see which model wrote the code
Knowing the source biases the score. Strip model names, agent logs, and formatting tells before review, and randomize the order patches appear in.
Using reviewers who do not know the stack
A general reviewer will catch syntax issues and miss a race condition in async Python or a subtle type hole in TypeScript. Match reviewers to the language and domain of the task.
Never refreshing the task set
If you reuse the same tasks for months, prompts and models get tuned to them. Pull fresh tasks from real tickets and production failures. Our piece on building an eval set from production failures walks through that sourcing step.
Where the rubric fits in your release process
A rubric is most useful when it runs at the same points every time: before you switch models, before you ship a prompt or agent change, and on a sample of real traffic each month. Pair it with automated checks so humans only spend time on what tests cannot see. For the automated side, see how to catch regressions in CI when models update, and how to write acceptance criteria for an AI feature so the brief your rubric checks against is clear in the first place.
If you would rather not build the reviewer bench yourself, Boundev runs expert code and patch evaluation for AI teams, with qualified reviewers, rubric scores, and written rationales for every judgment.
Frequently asked questions
How many criteria should a code grading rubric have?
Three to five. Correctness, security, and maintainability cover most needs. Add scope discipline and test quality when you evaluate agents that edit multiple files or write their own tests.
Should I use a 1 to 5 scale or pass/fail?
Use pass/fail for hard requirements such as security and "solves the task". Use a 1 to 5 scale with written anchors for qualities that vary by degree, such as maintainability.
How many patches do I need to validate a rubric?
Thirty to fifty independently double-scored patches is enough to find vague anchors and get a first agreement estimate. Grow the set as you find new failure types.
Can an LLM judge replace human reviewers?
For narrow, well-anchored criteria it can handle much of the volume, but only after you calibrate it against human scores on the same patches. Keep humans on security, borderline cases, and a regular audit sample.
Put expert judgment to work.
Talk to our team about evaluating AI-generated code, comparing responses, and building clear rubrics.