THE EVALUATION NOTEBOOK
A practical guide to AI evaluation
Worked examples, review rubrics, and practical guides for teams evaluating AI-generated software.
SOFTWARE EVALUATION PILOT
A focused pilot for your next evaluation decision.
Bring one question about your AI’s code or responses. Together, we’ll define the artifacts, review criteria, and evidence you need.
ILLUSTRATIVE REPORT · VERSION 1.0
A review you can trace, from requirement to reason.
A worked example of session-validation review, with original code, explicit requirements, and downloadable records.
EVALUATION BASICS
What passing tests can leave unanswered.
Use test results alongside explicit requirements and software judgment when reviewing AI-generated changes.
REVIEW QUALITY
Make reviewer disagreement useful.
A disagreement can reveal ambiguous requirements, weak score anchors, or missing evidence. Preserve enough detail to tell which.
RUBRIC DESIGN
Choose review dimensions that answer your question.
A compact rubric should make a decision easier, with explicit anchors and a way to handle missing evidence.
EVALUATION DESIGN
Build a sample that matches the decision.
Make the population, selection method, and exclusions explicit before interpreting an evaluation result.
EVALUATION WORKFLOW
Combine automated checks with expert review.
Give each form of evidence a clear job, and keep the reasoning behind the final decision visible.
REVIEW OPERATIONS
A practical adjudication playbook.
Define escalation rules before review begins, then preserve the evidence behind each final outcome.
YOUR NEXT EVALUATION
Start with the decision.
Share a short brief. Scope, availability, and fees are agreed before work starts.
