A month on the feedback loop
What changed when we forced every accepted task to carry a rubric note back to the expert. Spoiler: dispute rate dropped, tier progression accelerated.
Written by
Lukas Brennan
Principal Researcher
A month ago we made one change: every accepted task carries a one-paragraph rubric note back to the expert who produced it. Not as an option. Not as a 'reviewer comments' field that sits empty on the good ones. As a required field, on every accepted submission, written by the reviewer in plain language.
The results in a single month surprised us, including the parts that did not change.
What dropped
Dispute rate fell 38% across the network and 51% in the cohort of experts in their first ninety days. Most disputes were never really about the rejection — they were about the silence around the rejection. Once experts could see what specifically was being noted on accepted work, the rejections stopped feeling arbitrary.
- Dispute rate: −38% network-wide, −51% for experts in onboarding.
- Median time-to-tier-up: 11 weeks → 7 weeks.
- Reviewer time per item: +90 seconds. Worth it.
- Expert NPS: +14, with the largest gains in the Practitioner tier.
“The reviewer note on accepted work was doing more onboarding than our onboarding was.”
What did not change
First-pass acceptance rate did not move. We had quietly hoped it would, because better feedback should mean better next-batch submissions. It is possible we will see it in month two or three. It is also possible that acceptance rate is set by rubric design, not by individual feedback — and that is a different lever to pull.
What we are doing next
We are A/B testing structured notes (a short rubric-aligned tag plus the free-text paragraph) against pure free-text. We are also opening the notes corpus to a small group of senior arbiters, who are using it to find rubric clauses that produce inconsistent commentary — a leading indicator that the clause itself is ambiguous and due for a rewrite.
We will share the three-month read in due course. For now: the cheapest improvement we made all year was a required text field.
Principal Researcher
Lukas Brennan
Lukas studies how expert disagreement gets surfaced, calibrated, and turned into training signal.
More from the network on the same questions.
What rubrics can — and cannot — do
Rubrics are how we make judgement teachable. They are also the easiest thing to over-engineer. A short essay on the line between calibration and bureaucracy.
Training data and the judgement gap
Why the next leap in model quality will not come from more tokens, but from more disagreement — and how to elicit the right kind of disagreement from experts.
Have something to say?
Pitch us an essay
Guest posts from Lona experts and serious practitioners are welcome.