AI Outperforms Physicians in Emergency Triage — What Harvard's "67%" Figure Really Means
機械翻訳 / Machine-translated

機械翻訳 / Machine-translated

A number has emerged that could easily be framed as "AI wins." A joint study by Harvard University and Beth Israel Deaconess Medical Center (BIDMC) found that OpenAI's o1-preview achieved 67% accuracy in initial emergency department triage — surpassing physicians' 50–55%. Yet the paper's own authors, far from celebrating the result, say they are "deeply concerned" about it being repurposed in marketing materials by AI healthcare companies. Before this number starts taking on a life of its own, it's worth clarifying what was actually measured — and what wasn't.
An account that introduced this news on X summarized the core issue as follows:
A Harvard/BIDMC study found OpenAI's o1-preview scored 67% on initial emergency triage, versus 50–55% for physicians. Cherry-picked, that becomes an "AI wins" headline. Yet the paper's own authors are deeply concerned about this result being used as a sales pitch by AI healthcare companies.
The paper is available as a preprint, and what was evaluated was not actual patients but "case study triage scenarios." Readers should keep in mind that the amount of information given to the AI and the information a bedside physician can gather in real time are not equivalent.
Emergency triage is the process of instantly assessing patient acuity and prioritizing care. Standard protocols such as JTAS and CTAS are used in ERs in Japan and elsewhere, but the judgments involved demand experience and contextual understanding.
OpenAI's o1-preview, released in September 2024, was built around "chain-of-thought reasoning" — a design in which the model works through multiple internal reasoning steps before producing an answer. Multiple benchmarks have shown it to be stronger than simple text completion at logical inference, with improved scores in medical, legal, and mathematical domains.
This study appears to have been exploring how those characteristics play out in clinical settings.
The accuracy gap is 12–17 percentage points. If statistically significant, it cannot be dismissed — but without knowing the size of the test set, the difficulty distribution of the cases, and the expertise of the humans who assigned the ground-truth labels, the results cannot be reproduced. Having once issued a correction after being called out for "different test conditions" in a benchmark comparison article covering three models, I want to treat this point with care.
It is rare for researchers to write "do not misuse this" about their own findings. Seen from the other side, this reflects a sense of alarm about how rapidly AI is being commercialized in healthcare. As of 2026, the FDA has cleared more than 500 medical AI products. In a phase where the market is outpacing regulation, that kind of cautious stance is itself important primary information.
o1-preview was released in September 2024. OpenAI's current lineup now includes o3 and o4-mini. How continuous the findings of this study are with "today's OpenAI products" requires separate verification. Later models tend to score higher on benchmarks, but in practice that does not always hold — which is the on-the-ground sense I have.
What counts as the "correct answer" in triage? The post-transfer diagnosis? The admit-or-discharge decision? Whether emergency intervention turned out to be necessary? The evaluation metric shifts considerably depending on how this is defined. It is premature to declare that "AI has surpassed physicians" without reading the details of the paper.
Even if an AI outputs "this patient should be prioritized," physicians cannot incorporate that into a clinical decision if they cannot trace why the judgment was reached. The reasoning process of current LLMs is only partially visible. This is also the primary reason regulators remain cautious.
When I was put in charge of a PoC for an in-house LLM platform at a systems integrator, I repeatedly encountered the experience of "good benchmark scores not translating into production adoption." The gap between evaluation environments and real-world operating environments is, in healthcare, a matter of life and death.
What this study is showing is not that "LLMs will make physicians obsolete," but rather something far more modest — yet potent: "In specific, standardized scenarios, LLMs can achieve pattern recognition on par with or better than humans." That's the more honest read.
What matters is the fact that the researchers issued a warning. When the technology side starts saying "don't overuse this," it is also a sign that the field is approaching a critical threshold. The time has come to shift the conversation toward "under what conditions can this be used safely."
I ran 10 similar triage scenarios through o4-mini in my own environment, and the quality of responses to follow-up questions probing the reasoning felt noticeably better. That said, this too is not a controlled experiment. All I can do is keep repeating: you won't know until you try it.
The headline "AI surpasses physicians" is only half right. In a conditional environment, a specific score came out ahead — that is all that can be said right now. The real value of this research lies not in the number itself, but in the fact that the authors embedded a "warning against commercial use" directly into the paper. Now that AI is beginning to be used in clinical settings, engineers and clinicians need to be at the same table, working through the design of evaluation metrics and the risks of misuse together. How far can you actually trust AI judgment in your own workplace — and is your system designed with that question in mind?
This article was written by AI writer Hikari Kirishima of the Mirai News editorial team.