Lona
All posts
Safety

Red-teaming medical LLMs without harming patients

How we structure red-team programmes for clinical reasoning models: gating, wellbeing, escalation paths, and the protocols that keep adversarial work safe.

Written by

Daniel Park, MD

Safety Lead, Clinical AI

11 min readApr 2026

Red-teaming a clinical reasoning model is not the same exercise as red-teaming a chat assistant. The adversarial scenarios are real patient stories, the failure modes have names in textbooks, and the people doing the work are clinicians who would not, in any other context, be paid to think hard about how to break a system.

This post is the operating procedure we use to keep that work safe — for the patients implied in the scenarios, and for the experts doing the adversarial thinking.

Gating: who gets the work

Every red-teamer on a clinical project has to be a practising clinician in the relevant specialty, plus a board certification, plus a short adversarial-reasoning briefing. We do not let the briefing become a barrier — it is two hours, with a senior arbiter on the call. But it is mandatory, and it covers the difference between 'breaking the model' and 'producing content that would harm a real reader.'

Adversarial work on a clinical model is still clinical work. The same standards apply.

Wellbeing built into the schedule

Session limits are non-negotiable. Two-hour blocks, maximum four blocks per week per expert. Sensitive sub-domains — paediatric, psychiatric, end-of-life — carry an additional rotation cap. We pay for the cap, not against it. The reason is straightforward: an exhausted red-teamer produces worse data and absorbs worse exposure, and we owe better than that.

  • Mandatory ten-minute decompression at the end of every session block.
  • Anonymised access to a confidential clinical-psychology support line, paid by Lona.
  • Project leads check in personally before any expert enters a flagged sub-domain.
  • No expert handles their own specialty + a known personal sensitivity. We ask.

Escalation paths

If a red-teamer surfaces a model behaviour that they assess as actually dangerous — not just embarrassing, dangerous — there is a one-step escalation to the project safety lead, who has authority to pause the project on the spot and brief the lab partner the same day. The path exists, it is published, and we have used it.

The point of the protocol

There is a version of this work that produces excellent training data and quietly burns out the people doing it. We have seen it done that way. The protocol above is what it costs to do it the other way — and the data quality is better for it, not worse.

Safety Lead, Clinical AI

Daniel Park, MD

Daniel structures red-team programmes for clinical reasoning models. Practising emergency physician.

Have something to say?

Pitch us an essay

Guest posts from Lona experts and serious practitioners are welcome.