Datos de entrenamiento y la brecha de juicio
Por qué el próximo salto en la calidad del modelo no vendrá de más tokens, sino de más desacuerdo — y cómo obtener el tipo correcto de desacuerdo de los expertos.
Escrito por
Mira Okafor
Head of Quality
The cheap way to make a model better used to be: feed it more tokens. That curve is bending. The next leap in capability will not come from another order of magnitude of text — it will come from teaching models to handle the small, hard cases that careful experts disagree about.
We call that gap between what a model can pattern-match and what a human expert would actually do the judgement gap. It is invisible in benchmarks that average across thousands of easy questions. It shows up the moment a model has to triage, prioritise, or refuse.
Disagreement is the signal
When three senior radiologists look at the same scan and give different reads, that is not noise. That is the model's curriculum. The interesting training data is exactly where well-calibrated experts split — and where, after a short discussion, they converge on a defensible answer.
Most data pipelines are built to suppress that signal. They aggregate to majority vote, throw out the dissent, and label the rest 'gold.' The result is a model that learns the average opinion and never the reasoning that made the minority opinion defensible.
“We were not collecting answers. We were collecting the disagreements we wished we had heard the first time.”
Eliciting the right kind of disagreement
There is bad disagreement — two annotators reading the rubric differently, or one of them rushing through a batch. Then there is good disagreement — two experts who fully understand the question and who reach different defensible answers because their training, their patient population, or their risk tolerance diverges.
Our job is to design tasks that surface the second kind and filter out the first. Three things make that work in practice:
- A rubric that makes the implicit explicit — what to optimise for, what to never trade away, and what is acceptable variance.
- Independent first passes — no anchoring, no shared chat threads until each expert has committed to a position.
- A structured reconciliation pass where each expert reads the other reads, then either updates or doubles down with a one-paragraph rationale.
What we are not claiming
We are not claiming more disagreement is always better. We are claiming the right disagreement, made legible by a rubric and a reconciliation step, is the most efficient training signal we have found for the cases that matter most. The judgement gap closes one well-structured argument at a time.
Head of Quality
Mira Okafor
Mira leads quality and rubric design at Lona. Previously: senior research scientist at two frontier labs.
Más de la red sobre las mismas cuestiones.
Lo que las rúbricas pueden — y no pueden — hacer
Las rúbricas son cómo hacemos que el juicio sea enseñable. También son lo más fácil de sobre-diseñar. Un breve ensayo sobre la línea entre la calibración y la burocracia.
El caso por la experiencia vertical en el entrenamiento de IA
Los anotadores generalistas tienen un límite. Abogamos por una división más clara: especialistas en la materia para el contenido, lingüistas para el tono, revisores calibrados para ambos.
¿Tienes algo que decir?
Propón un ensayo
Se aceptan publicaciones de invitados de expertos de Lona y profesionales serios.