Lona
Todas las publicaciones
Calidad

Datos de entrenamiento y la brecha de juicio

Por qué el próximo salto en la calidad del modelo no vendrá de más tokens, sino de más desacuerdo — y cómo obtener el tipo correcto de desacuerdo de los expertos.

Escrito por

Mira Okafor

Head of Quality

7 min de lecturaMayo 2026

The cheap way to make a model better used to be: feed it more tokens. That curve is bending. The next leap in capability will not come from another order of magnitude of text — it will come from teaching models to handle the small, hard cases that careful experts disagree about.

We call that gap between what a model can pattern-match and what a human expert would actually do the judgement gap. It is invisible in benchmarks that average across thousands of easy questions. It shows up the moment a model has to triage, prioritise, or refuse.

Disagreement is the signal

When three senior radiologists look at the same scan and give different reads, that is not noise. That is the model's curriculum. The interesting training data is exactly where well-calibrated experts split — and where, after a short discussion, they converge on a defensible answer.

Most data pipelines are built to suppress that signal. They aggregate to majority vote, throw out the dissent, and label the rest 'gold.' The result is a model that learns the average opinion and never the reasoning that made the minority opinion defensible.

We were not collecting answers. We were collecting the disagreements we wished we had heard the first time.
Mira Okafor

Eliciting the right kind of disagreement

There is bad disagreement — two annotators reading the rubric differently, or one of them rushing through a batch. Then there is good disagreement — two experts who fully understand the question and who reach different defensible answers because their training, their patient population, or their risk tolerance diverges.

Our job is to design tasks that surface the second kind and filter out the first. Three things make that work in practice:

  • A rubric that makes the implicit explicit — what to optimise for, what to never trade away, and what is acceptable variance.
  • Independent first passes — no anchoring, no shared chat threads until each expert has committed to a position.
  • A structured reconciliation pass where each expert reads the other reads, then either updates or doubles down with a one-paragraph rationale.

What we are not claiming

We are not claiming more disagreement is always better. We are claiming the right disagreement, made legible by a rubric and a reconciliation step, is the most efficient training signal we have found for the cases that matter most. The judgement gap closes one well-structured argument at a time.

Head of Quality

Mira Okafor

Mira leads quality and rubric design at Lona. Previously: senior research scientist at two frontier labs.

¿Tienes algo que decir?

Propón un ensayo

Se aceptan publicaciones de invitados de expertos de Lona y profesionales serios.