OpenAI's o1 preview model out-diagnosed emergency room physicians in a head-to-head Harvard Medical School study, raising fresh questions about how quickly large language models could move into frontline triage. Researchers at Beth Israel Deaconess fed the model the same electronic medical records, vitals and intake nurse notes given to two attending internal medicine doctors across 76 real ER cases.
At the triage stage, where clinicians have the least information and the least time, o1 landed on the exact or near-exact diagnosis 67% of the time. The two attending physicians scored 55% and 50%. "The O1 model hit the exact or very close diagnosis 67% of the time. And the two attending, the two attending doctors hit it 55 to 50% of the time," Jaeden Schafer said on the podcast.
Accuracy rose for everyone once patients were admitted and more data became available, but the model still led. o1 reached 81.6%, while the two doctors scored 78.9% and 69.7%. The gap was widest precisely where ER medicine is hardest: the opening minutes of an unfamiliar case.
Key facts
- 01OpenAI's o1 preview correctly diagnosed 67% of 76 real ER triage cases, versus 50–55% for two attending internal medicine physicians.
- 02Once patients were admitted, o1 hit 81.6% diagnostic accuracy, ahead of the two doctors at 78.9% and 69.7%.
- 03The Harvard Medical School and Beth Israel Deaconess team ran six experiments pitting the model against hundreds of physicians.
- 04The study used text-only inputs — no imaging or multimodal data — and the authors stop short of endorsing clinical deployment.
The scope of the work goes beyond a single matchup. The team ran six experiments and tested the model against hundreds of physicians across a range of cases. "They did this across six total experiments and the model went up against hundreds of different physicians and it kind of held its own in one," Schafer said, pushing back on the assumption that the result rests on a small or cherry-picked sample.
“The O1 model hit the exact or very close diagnosis 67% of the time. And the two attending, the two attending doctors hit it 55 to 50% of the time”— Jaeden Schafer
Peter Broder, a Harvard clinical fellow at Beth Israel and one of the study's co-authors, told Fortune that traditional evaluation methods have run out of road, noting that models are already scoring close to 100% on multiple choice tests. The implication is that medical AI benchmarks need to move toward the messier, real-world workflows the Harvard team simulated.
Important caveats remain. The inputs were text-only, with no medical imaging or other multimodal signals, and the authors are explicit that the work is not an endorsement of clinical deployment. ER medicine is also a specific slice of practice — fast, anonymous, high-stakes — and unlikely to mirror the gains a model might or might not deliver in primary care, where physicians know their patients' histories.
The thornier issue is legal. Hospitals adopting tools like o1 will quickly bump into an unresolved liability question about who carries the risk when human and machine disagree. "If O1 disagrees with the attending physicians and then the attendings override it, who's going to get sued when a patient has an issue for malpractice?" Schafer asked, framing a question malpractice insurers and hospital systems have yet to answer.
For OpenAI, the result is another data point in its push to position frontier reasoning models as serious tools for high-stakes professional work, not just chat. For emergency medicine, it is an early signal that triage — long considered one of the hardest judgment calls in the hospital — may be where AI lands first, well before regulators, insurers and courts have decided what to do about it.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




