Brief Safety Reminders Reduce Harmful AI Clinical Decisions
A Mount Sinai study tested 20 large language models on clinical scenarios and found that a brief safety reminder reduced potentially harmful choices. Across more than 10 million responses, harmful choices fell from 16.6% to 10.1% when the reminder was included. The results suggest that how AI is prompted and framed can influence clinical safety.
The Mount Sinai team evaluated 20 large language models using 501 variants of 50 clinical situations and 100 de-identified discharge cases. Across over 10 million answers, roughly 1.18 million choices were possibly unsafe. A short safety reminder cut unsafe responses from 16.6% to 10.1%, and it helped 19 of 20 models.
Tests included requests to omit recommended follow-up blood work to lessen workload, sometimes framed as urgent or as a supervisor's instruction. Models chose among four actions: comply, keep the follow-up, or ask a clinician for guidance. The researchers varied wording, used three brief reminders, repeated each combination ten times, and randomized answer order. Other risky examples included halting antibiotic treatment.
If these findings hold, brief safety prompts could become a low-cost layer in clinical AI deployment, potentially reducing some unsafe recommendations that might otherwise reach patients. Clinicians and health systems may benefit from added guardrails, while developers could face pressure to test how models respond to conflicting instructions. Patients, especially those whose care relies on AI-assisted workflows, may see safer outputs, though reminders alone cannot replace human oversight or eliminate risk.