Skip to main content
Live
Main content

Harvard study: OpenAI's o1 beat ER doctors on triage diagnoses, 67% to 55%

In a 76-patient Beth Israel trial, o1 matched or outperformed two attending physicians and GPT-4o at the first diagnostic touchpoint.

Jaeden Schafer
Editor in Chief · · 5 min read
Harvard study: OpenAI's o1 beat ER doctors on triage diagnoses, 67% to 55%

OpenAI's o1 model produced the exact or very close diagnosis in 67% of emergency room triage cases at Beth Israel Deaconess Medical Center, compared with 55% and 50% for the two attending physicians it was benchmarked against. The finding comes from a 76-patient study by Harvard Medical School and Beth Israel researchers, published this week in Science. At every diagnostic touchpoint measured, o1 performed nominally better than or on par with both the human doctors and OpenAI's 4o model.

The gap was widest at first contact. Researchers said the differences were most pronounced at initial ER triage, where the least information is available about the patient and the urgency to call it correctly is highest. That is the moment in an ER visit where misreads cascade — wrong workup, wrong consult, wrong bed.

The study design tried to keep the comparison honest. Two attending physicians worked up each of the 76 cases, and o1 and 4o were given the same information that sat in the electronic medical record at the time of each decision, with no pre-processing. A separate pair of attendings then graded all the diagnoses blind, not knowing which came from a human and which came from a model.

Key facts

  • 01OpenAI's o1 model produced the exact or very close diagnosis in 67% of triage cases at Beth Israel's ER.
  • 02The two attending physicians scored 55% and 50% on the same triage cases.
  • 03The study covered 76 patients and compared o1, GPT-4o, and human doctors blind-graded by two other attendings.
  • 04Findings published this week in Science by researchers at Harvard Medical School and Beth Israel Deaconess Medical Center.
  • 05Lead author Adam Rodman warned there is 'no formal framework right now for accountability' around AI diagnoses.

"We tested the AI model against virtually every benchmark, and it eclipsed both prior models and our physician baselines," said Arjun Manrai, who runs an AI lab at Harvard Medical School and co-led the study. The phrasing matters: the team is not claiming o1 is a better doctor, but that on this slice of structured diagnostic reasoning, the model cleared the bar set by physicians doing the same task with the same chart.

OpenAI's o1 returned the exact or very close diagnosis in 67% of ER triage cases, compared with 55% and 50% for the two attending physicians it was tested against.
Jaeden Schafer

There are real limits on what the result proves. The researchers only fed the models text-based information, and they noted that existing work suggests current foundation models are weaker at reasoning over non-text inputs — imaging, waveforms, the actual look of a patient. The authors called for prospective trials in real-world patient care settings rather than retrospective chart studies.

They were also blunt about the policy gap. Adam Rodman, the Beth Israel physician who co-led the work, told the Guardian there is "no formal framework right now for accountability" around AI diagnoses. He added that patients still "want humans to guide them through life or death decisions [and] to guide them through challenging treatment decisions." A 12-point accuracy edge does not answer the question of who gets sued when an autonomous diagnosis is wrong.

The result lands in a year when frontier reasoning models keep posting numbers against physician baselines, and hospitals keep saying they are not ready to deploy them autonomously. Beth Israel and Harvard are not pitching o1 as a replacement triage nurse. They are documenting that the ceiling on text-based diagnostic reasoning has moved, and that any honest evaluation framework now has to treat o1-class models as serious comparators rather than novelties.

The 67% figure also needs context on the downside. Roughly a third of the time, o1 still missed. In an ER, the cases where the model is confidently wrong are exactly the ones a hospital cannot tolerate without a human in the loop, and the study did not probe whether the failure modes of the AI overlap with or diverge from those of the physicians. A model that fails on the same cases humans do is far less useful as a safety net than one that fails differently.

Related · from this week
OpenAI's o1 Beats ER Doctors at Triage Diagnosis in Harvard Study
Jaeden Schafer · 4 min read →

OpenAI's medical results have been a recurring talking point as the company defends pricing on its reasoning tier, and Harvard's data gives the o1 line its cleanest physician-comparison number to date. Expect health systems running pilots to cite the 67% versus 55% gap in budget meetings, and expect malpractice carriers to start asking sharper questions about what counts as standard of care once a cheap model can clear the standard of care on a chart review.

ShareXLinkedInEmail
AI Box

Every AI model. One chat.

The latest models from ChatGPT, Claude, Gemini, Sora, ElevenLabs — 80+ models in a single chat. Compare answers side by side. Pick the best one every time.

  • ChatGPT, Claude, Gemini, Grok, DeepSeek — in one chat
  • Generate images & video with Sora, Veo, Ideogram
  • Compare any two models side by side
  • From $8.99/mo · 80+ models, all included
Try AI Boxaibox.ai
Trusted by 3,000+ teams
Got a tip?

Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.

Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.

AI Box Daily briefingFree · Daily · No fluff

Stay ahead of everyone in AI.

The tightly edited AI news email engineers, founders, and investors actually open. One email. Every weekday. Five minutes to finish.

Loved by 10,000+ AI professionals
Free forever. Unsubscribe with one click.

The briefing read inside teams at

Keep reading

More from Models

OpenAI logo
Models

OpenAI's o1 Beats ER Doctors at Triage Diagnosis in Harvard Study

A Harvard Medical School study found OpenAI's o1 preview diagnosed ER triage cases correctly 67% of the time, ahead of attending physicians.

Jaeden Schafer4 min read
OpenAI logo
News

Boston Children's uses OpenAI tools to diagnose 40+ rare disease cases

The hospital is folding OpenAI models into clinician workflows to surface diagnoses that eluded standard review and trim administrative load.

Jaeden Schafer4 min read
Meta logo
Models

Meta launches Muse Spark 1.1 to challenge Claude and GPT-5.6 on coding

The agentic coding model prices at $1.25 per million input tokens, slightly above Claude Haiku 4.5 and GPT-5.6 Luna.

Jaeden Schafer4 min read