Training data and the judgement gap
Why the next leap in model quality will not come from more tokens, but from more disagreement — and how to elicit the right kind of disagreement from experts.
Written by
Mira Okafor
Head of Quality
The cheap way to make a model better used to be: feed it more tokens. That curve is bending. The next leap in capability will not come from another order of magnitude of text — it will come from teaching models to handle the small, hard cases that careful experts disagree about.
We call that gap between what a model can pattern-match and what a human expert would actually do the judgement gap. It is invisible in benchmarks that average across thousands of easy questions. It shows up the moment a model has to triage, prioritise, or refuse.
Disagreement is the signal
When three senior radiologists look at the same scan and give different reads, that is not noise. That is the model's curriculum. The interesting training data is exactly where well-calibrated experts split — and where, after a short discussion, they converge on a defensible answer.
Most data pipelines are built to suppress that signal. They aggregate to majority vote, throw out the dissent, and label the rest 'gold.' The result is a model that learns the average opinion and never the reasoning that made the minority opinion defensible.
“We were not collecting answers. We were collecting the disagreements we wished we had heard the first time.”
Eliciting the right kind of disagreement
There is bad disagreement — two annotators reading the rubric differently, or one of them rushing through a batch. Then there is good disagreement — two experts who fully understand the question and who reach different defensible answers because their training, their patient population, or their risk tolerance diverges.
Our job is to design tasks that surface the second kind and filter out the first. Three things make that work in practice:
- A rubric that makes the implicit explicit — what to optimise for, what to never trade away, and what is acceptable variance.
- Independent first passes — no anchoring, no shared chat threads until each expert has committed to a position.
- A structured reconciliation pass where each expert reads the other reads, then either updates or doubles down with a one-paragraph rationale.
What we are not claiming
We are not claiming more disagreement is always better. We are claiming the right disagreement, made legible by a rubric and a reconciliation step, is the most efficient training signal we have found for the cases that matter most. The judgement gap closes one well-structured argument at a time.
Head of Quality
Mira Okafor
Mira leads quality and rubric design at Lona. Previously: senior research scientist at two frontier labs.
More from the network on the same questions.
What rubrics can — and cannot — do
Rubrics are how we make judgement teachable. They are also the easiest thing to over-engineer. A short essay on the line between calibration and bureaucracy.
The case for vertical expertise in AI training
Generalist annotators have a ceiling. We argue for a sharper division: subject-matter specialists for substance, linguists for tone, calibrated reviewers for both.
Have something to say?
Pitch us an essay
Guest posts from Lona experts and serious practitioners are welcome.