OpenAI's o1 model produced the exact or very close diagnosis in 67% of emergency room triage cases at Beth Israel Deaconess Medical Center, compared with 55% and 50% for the two attending physicians it was benchmarked against. The finding comes from a 76-patient study by Harvard Medical School and Beth Israel researchers, published this week in Science. At every diagnostic touchpoint measured, o1 performed nominally better than or on par with both the human doctors and OpenAI's 4o model.
The gap was widest at first contact. Researchers said the differences were most pronounced at initial ER triage, where the least information is available about the patient and the urgency to call it correctly is highest. That is the moment in an ER visit where misreads cascade — wrong workup, wrong consult, wrong bed.
The study design tried to keep the comparison honest. Two attending physicians worked up each of the 76 cases, and o1 and 4o were given the same information that sat in the electronic medical record at the time of each decision, with no pre-processing. A separate pair of attendings then graded all the diagnoses blind, not knowing which came from a human and which came from a model.
Key facts
- 01OpenAI's o1 model produced the exact or very close diagnosis in 67% of triage cases at Beth Israel's ER.
- 02The two attending physicians scored 55% and 50% on the same triage cases.
- 03The study covered 76 patients and compared o1, GPT-4o, and human doctors blind-graded by two other attendings.
- 04Findings published this week in Science by researchers at Harvard Medical School and Beth Israel Deaconess Medical Center.
- 05Lead author Adam Rodman warned there is 'no formal framework right now for accountability' around AI diagnoses.
"We tested the AI model against virtually every benchmark, and it eclipsed both prior models and our physician baselines," said Arjun Manrai, who runs an AI lab at Harvard Medical School and co-led the study. The phrasing matters: the team is not claiming o1 is a better doctor, but that on this slice of structured diagnostic reasoning, the model cleared the bar set by physicians doing the same task with the same chart.
“OpenAI's o1 returned the exact or very close diagnosis in 67% of ER triage cases, compared with 55% and 50% for the two attending physicians it was tested against.”— Jaeden Schafer
There are real limits on what the result proves. The researchers only fed the models text-based information, and they noted that existing work suggests current foundation models are weaker at reasoning over non-text inputs — imaging, waveforms, the actual look of a patient. The authors called for prospective trials in real-world patient care settings rather than retrospective chart studies.
They were also blunt about the policy gap. Adam Rodman, the Beth Israel physician who co-led the work, told the Guardian there is "no formal framework right now for accountability" around AI diagnoses. He added that patients still "want humans to guide them through life or death decisions [and] to guide them through challenging treatment decisions." A 12-point accuracy edge does not answer the question of who gets sued when an autonomous diagnosis is wrong.
The result lands in a year when frontier reasoning models keep posting numbers against physician baselines, and hospitals keep saying they are not ready to deploy them autonomously. Beth Israel and Harvard are not pitching o1 as a replacement triage nurse. They are documenting that the ceiling on text-based diagnostic reasoning has moved, and that any honest evaluation framework now has to treat o1-class models as serious comparators rather than novelties.
The 67% figure also needs context on the downside. Roughly a third of the time, o1 still missed. In an ER, the cases where the model is confidently wrong are exactly the ones a hospital cannot tolerate without a human in the loop, and the study did not probe whether the failure modes of the AI overlap with or diverge from those of the physicians. A model that fails on the same cases humans do is far less useful as a safety net than one that fails differently.
OpenAI's medical results have been a recurring talking point as the company defends pricing on its reasoning tier, and Harvard's data gives the o1 line its cleanest physician-comparison number to date. Expect health systems running pilots to cite the 67% versus 55% gap in budget meetings, and expect malpractice carriers to start asking sharper questions about what counts as standard of care once a cheap model can clear the standard of care on a chart review.
Working on something we should cover, or seeing a story we missed? Send leads, documents, or feedback to hello@aichatdaily.com. For sensitive tips, see our secure tips page for Signal and PGP options.
Spotted an error? Email hello@aichatdaily.com with the URL and the issue, or read our full corrections policy.




