Lo que las rúbricas pueden — y no pueden — hacer
Las rúbricas son cómo hacemos que el juicio sea enseñable. También son lo más fácil de sobre-diseñar. Un breve ensayo sobre la línea entre la calibración y la burocracia.
Escrito por
Lukas Brennan
Principal Researcher
A good rubric is one of the highest-leverage artefacts in a training pipeline. It is also the easiest thing in the building to over-engineer. The same instinct that produces a great rubric — be precise, be explicit, leave nothing to interpretation — produces a terrible one if you keep pulling on it.
Rubrics make judgement teachable. They do not, and cannot, replace it.
What a rubric is for
Three concrete jobs: align experts before they start, give reviewers a stable basis for accepting or rejecting work, and let new people get to competent output without years of apprenticeship. When a rubric does those three things, you stop having the same argument every Friday.
A rubric that is doing its job is short, opinionated, and has worked examples next to anti-examples. It says what to optimise for, what is acceptable variance, and which trade-offs are off the table. It does not try to be a complete decision procedure.
“If your rubric is longer than the task it governs, you have built a different task.”
The over-engineering trap
Every edge case wants to become a rule. Every reviewer disagreement wants to become a clause. After six months of unchecked accretion, the rubric is forty pages long, internally contradictory, and no longer faster to read than to ignore. People stop reading it and start guessing what the reviewers want — which is exactly the state you wrote the rubric to escape.
- Hard cap: one page of rubric for a task that takes under thirty minutes.
- Every clause earns its place by pointing at a real, recurring disagreement, not a hypothetical one.
- Worked example next to every rule. If you cannot show the rule in action, the rule is not ready.
- Quarterly prune: rules that have not been cited in a review get deleted.
Calibration over compliance
The goal of a rubric is not to make every expert produce identical output. It is to make the differences interpretable. Two experts following a good rubric should still disagree, sometimes — but their disagreements should be legible, and the reconciliation should be quick. That is what calibration looks like in practice, and it is the only place where rubric work actually pays off.
Principal Researcher
Lukas Brennan
Lukas studies how expert disagreement gets surfaced, calibrated, and turned into training signal.
Más de la red sobre las mismas cuestiones.
Datos de entrenamiento y la brecha de juicio
Por qué el próximo salto en la calidad del modelo no vendrá de más tokens, sino de más desacuerdo — y cómo obtener el tipo correcto de desacuerdo de los expertos.
Un mes en el bucle de retroalimentación
Qué cambió cuando obligamos a cada tarea aceptada a llevar una nota de rúbrica de vuelta al experto. Spoiler: la tasa de disputas disminuyó, la progresión de niveles se aceleró.
¿Tienes algo que decir?
Propón un ensayo
Se aceptan publicaciones de invitados de expertos de Lona y profesionales serios.