What rubrics can — and cannot — do
Rubrics are how we make judgement teachable. They are also the easiest thing to over-engineer. A short essay on the line between calibration and bureaucracy.
Written by
Lukas Brennan
Principal Researcher
A good rubric is one of the highest-leverage artefacts in a training pipeline. It is also the easiest thing in the building to over-engineer. The same instinct that produces a great rubric — be precise, be explicit, leave nothing to interpretation — produces a terrible one if you keep pulling on it.
Rubrics make judgement teachable. They do not, and cannot, replace it.
What a rubric is for
Three concrete jobs: align experts before they start, give reviewers a stable basis for accepting or rejecting work, and let new people get to competent output without years of apprenticeship. When a rubric does those three things, you stop having the same argument every Friday.
A rubric that is doing its job is short, opinionated, and has worked examples next to anti-examples. It says what to optimise for, what is acceptable variance, and which trade-offs are off the table. It does not try to be a complete decision procedure.
“If your rubric is longer than the task it governs, you have built a different task.”
The over-engineering trap
Every edge case wants to become a rule. Every reviewer disagreement wants to become a clause. After six months of unchecked accretion, the rubric is forty pages long, internally contradictory, and no longer faster to read than to ignore. People stop reading it and start guessing what the reviewers want — which is exactly the state you wrote the rubric to escape.
- Hard cap: one page of rubric for a task that takes under thirty minutes.
- Every clause earns its place by pointing at a real, recurring disagreement, not a hypothetical one.
- Worked example next to every rule. If you cannot show the rule in action, the rule is not ready.
- Quarterly prune: rules that have not been cited in a review get deleted.
Calibration over compliance
The goal of a rubric is not to make every expert produce identical output. It is to make the differences interpretable. Two experts following a good rubric should still disagree, sometimes — but their disagreements should be legible, and the reconciliation should be quick. That is what calibration looks like in practice, and it is the only place where rubric work actually pays off.
Principal Researcher
Lukas Brennan
Lukas studies how expert disagreement gets surfaced, calibrated, and turned into training signal.
More from the network on the same questions.
Training data and the judgement gap
Why the next leap in model quality will not come from more tokens, but from more disagreement — and how to elicit the right kind of disagreement from experts.
A month on the feedback loop
What changed when we forced every accepted task to carry a rubric note back to the expert. Spoiler: dispute rate dropped, tier progression accelerated.
Have something to say?
Pitch us an essay
Guest posts from Lona experts and serious practitioners are welcome.