Equipos de ataque a LLMs médicos sin dañar a los pacientes
Cómo estructuramos programas de equipos de ataque para modelos de razonamiento clínico: control de acceso, bienestar, rutas de escalada y los protocolos que mantienen el trabajo adversarial seguro.
Escrito por
Daniel Park, MD
Safety Lead, Clinical AI
Red-teaming a clinical reasoning model is not the same exercise as red-teaming a chat assistant. The adversarial scenarios are real patient stories, the failure modes have names in textbooks, and the people doing the work are clinicians who would not, in any other context, be paid to think hard about how to break a system.
This post is the operating procedure we use to keep that work safe — for the patients implied in the scenarios, and for the experts doing the adversarial thinking.
Gating: who gets the work
Every red-teamer on a clinical project has to be a practising clinician in the relevant specialty, plus a board certification, plus a short adversarial-reasoning briefing. We do not let the briefing become a barrier — it is two hours, with a senior arbiter on the call. But it is mandatory, and it covers the difference between 'breaking the model' and 'producing content that would harm a real reader.'
“Adversarial work on a clinical model is still clinical work. The same standards apply.”
Wellbeing built into the schedule
Session limits are non-negotiable. Two-hour blocks, maximum four blocks per week per expert. Sensitive sub-domains — paediatric, psychiatric, end-of-life — carry an additional rotation cap. We pay for the cap, not against it. The reason is straightforward: an exhausted red-teamer produces worse data and absorbs worse exposure, and we owe better than that.
- Mandatory ten-minute decompression at the end of every session block.
- Anonymised access to a confidential clinical-psychology support line, paid by Lona.
- Project leads check in personally before any expert enters a flagged sub-domain.
- No expert handles their own specialty + a known personal sensitivity. We ask.
Escalation paths
If a red-teamer surfaces a model behaviour that they assess as actually dangerous — not just embarrassing, dangerous — there is a one-step escalation to the project safety lead, who has authority to pause the project on the spot and brief the lab partner the same day. The path exists, it is published, and we have used it.
The point of the protocol
There is a version of this work that produces excellent training data and quietly burns out the people doing it. We have seen it done that way. The protocol above is what it costs to do it the other way — and the data quality is better for it, not worse.
Safety Lead, Clinical AI
Daniel Park, MD
Daniel structures red-team programmes for clinical reasoning models. Practising emergency physician.
Más de la red sobre las mismas cuestiones.
Datos de entrenamiento y la brecha de juicio
Por qué el próximo salto en la calidad del modelo no vendrá de más tokens, sino de más desacuerdo — y cómo obtener el tipo correcto de desacuerdo de los expertos.
Lo que las rúbricas pueden — y no pueden — hacer
Las rúbricas son cómo hacemos que el juicio sea enseñable. También son lo más fácil de sobre-diseñar. Un breve ensayo sobre la línea entre la calibración y la burocracia.
¿Tienes algo que decir?
Propón un ensayo
Se aceptan publicaciones de invitados de expertos de Lona y profesionales serios.