Training data and the judgement gap
Why the next leap in model quality will not come from more tokens, but from more disagreement — and how to elicit the right kind of disagreement from experts.
Blog
Essays on training data, quality, process, and the practical work of building reliable AI.
Latest posts
One or two essays a week from our team and senior experts.
Why the next leap in model quality will not come from more tokens, but from more disagreement — and how to elicit the right kind of disagreement from experts.
Rubrics are how we make judgement teachable. They are also the easiest thing to over-engineer. A short essay on the line between calibration and bureaucracy.
A breakdown of how we set rates: market signals, project stakes, tier multipliers, and the deliberate decisions we made to avoid race-to-the-bottom dynamics.
How we structure red-team programmes for clinical reasoning models: gating, wellbeing, escalation paths, and the protocols that keep adversarial work safe.
Generalist annotators have a ceiling. We argue for a sharper division: subject-matter specialists for substance, linguists for tone, calibrated reviewers for both.
What changed when we forced every accepted task to carry a rubric note back to the expert. Spoiler: dispute rate dropped, tier progression accelerated.
Have something to say?
Guest posts from Lona experts and serious practitioners are welcome.