Skip to main content
Live
Main content

OpenAI's o1 Beats ER Doctors at Triage Diagnosis in Harvard Study

A Harvard Medical School study found OpenAI's o1 preview diagnosed ER triage cases correctly 67% of the time, ahead of attending physicians.

Jaeden Schafer
Editor in Chief · · 4 min read
OpenAI logo
AICD
AI Chat Podcast

Unreleased AI Models: Government's Interest

OpenAI's o1 preview model out-diagnosed emergency room physicians in a head-to-head Harvard Medical School study, raising fresh questions about how quickly large language models could move into frontline triage. Researchers at Beth Israel Deaconess fed the model the same electronic medical records, vitals and intake nurse notes given to two attending internal medicine doctors across 76 real ER cases.

At the triage stage, where clinicians have the least information and the least time, o1 landed on the exact or near-exact diagnosis 67% of the time. The two attending physicians scored 55% and 50%. "The O1 model hit the exact or very close diagnosis 67% of the time. And the two attending, the two attending doctors hit it 55 to 50% of the time," Jaeden Schafer said on the podcast.

Accuracy rose for everyone once patients were admitted and more data became available, but the model still led. o1 reached 81.6%, while the two doctors scored 78.9% and 69.7%. The gap was widest precisely where ER medicine is hardest: the opening minutes of an unfamiliar case.

Key facts

  • 01OpenAI's o1 preview correctly diagnosed 67% of 76 real ER triage cases, versus 50–55% for two attending internal medicine physicians.
  • 02Once patients were admitted, o1 hit 81.6% diagnostic accuracy, ahead of the two doctors at 78.9% and 69.7%.
  • 03The Harvard Medical School and Beth Israel Deaconess team ran six experiments pitting the model against hundreds of physicians.
  • 04The study used text-only inputs — no imaging or multimodal data — and the authors stop short of endorsing clinical deployment.

The scope of the work goes beyond a single matchup. The team ran six experiments and tested the model against hundreds of physicians across a range of cases. "They did this across six total experiments and the model went up against hundreds of different physicians and it kind of held its own in one," Schafer said, pushing back on the assumption that the result rests on a small or cherry-picked sample.

The O1 model hit the exact or very close diagnosis 67% of the time. And the two attending, the two attending doctors hit it 55 to 50% of the time
Jaeden Schafer

Peter Broder, a Harvard clinical fellow at Beth Israel and one of the study's co-authors, told Fortune that traditional evaluation methods have run out of road, noting that models are already scoring close to 100% on multiple choice tests. The implication is that medical AI benchmarks need to move toward the messier, real-world workflows the Harvard team simulated.

Important caveats remain. The inputs were text-only, with no medical imaging or other multimodal signals, and the authors are explicit that the work is not an endorsement of clinical deployment. ER medicine is also a specific slice of practice — fast, anonymous, high-stakes — and unlikely to mirror the gains a model might or might not deliver in primary care, where physicians know their patients' histories.

The thornier issue is legal. Hospitals adopting tools like o1 will quickly bump into an unresolved liability question about who carries the risk when human and machine disagree. "If O1 disagrees with the attending physicians and then the attendings override it, who's going to get sued when a patient has an issue for malpractice?" Schafer asked, framing a question malpractice insurers and hospital systems have yet to answer.

For OpenAI, the result is another data point in its push to position frontier reasoning models as serious tools for high-stakes professional work, not just chat. For emergency medicine, it is an early signal that triage — long considered one of the hardest judgment calls in the hospital — may be where AI lands first, well before regulators, insurers and courts have decided what to do about it.

Related · from this week
Harvard study: OpenAI's o1 beat ER doctors on triage diagnoses, 67% to 55%
Jaeden Schafer · 5 min read →
ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

Harvard study: OpenAI's o1 beat ER doctors on triage diagnoses, 67% to 55%
Models

Harvard study: OpenAI's o1 beat ER doctors on triage diagnoses, 67% to 55%

In a 76-patient Beth Israel trial, o1 matched or outperformed two attending physicians and GPT-4o at the first diagnostic touchpoint.

Jaeden Schafer5 min read
OpenAI logo
News

Boston Children's uses OpenAI tools to diagnose 40+ rare disease cases

The hospital is folding OpenAI models into clinician workflows to surface diagnoses that eluded standard review and trim administrative load.

Jaeden Schafer4 min read
Meta logo
Models

Meta launches Muse Spark 1.1 to challenge Claude and GPT-5.6 on coding

The agentic coding model prices at $1.25 per million input tokens, slightly above Claude Haiku 4.5 and GPT-5.6 Luna.

Jaeden Schafer4 min read