OpenAI o1 Outperforms Doctors in ER Triage — 67% vs. 55%, But Why the Authors Say "Don't Use It"
機械翻訳 / Machine-translated
A joint study by Harvard University and Beth Israel Deaconess Medical Center (BIDMC) has reported that OpenAI's o1-preview outperformed physicians in initial emergency department triage. Taken in isolation, the figures — 67% accuracy versus 50–55% — make for an easy "AI wins" headline. Yet the paper's own authors have publicly voiced concern about these results being used to promote medical AI companies — and that contradiction reveals the most important structure to understand today.
The setting is the "initial triage" stage of the emergency department — the process of classifying incoming patients into one of five urgency levels, where errors directly affect patient outcomes.
In an evaluation using real case data from BIDMC, o1-preview achieved a 67% accuracy rate, outperforming a group of experienced physicians by approximately 12 to 17 percentage points over their 50–55%.
"The paper's own authors are strongly concerned about this result being used as a selling point by AI healthcare companies."
— Summarizing the intent of a research introduction on X (paraphrased from the original)
As of May 2, 2026, the paper is circulating as a preprint, and it should be noted that it has not yet undergone peer-reviewed publication.
Research evaluating LLMs in the medical field surged after 2023, with a string of announcements foregrounding an "AI vs. doctor" narrative — including breakthroughs at the level of the United States Medical Licensing Examination (USMLE) and validations of GPT-4's diagnostic accuracy.
In clinical settings, however, there is a deep gap between "accuracy on a benchmark" and "safety in actual use." Triage is not a simple knowledge problem; it requires real-time integration of a patient's facial expression, vital signs, and interview responses — and accuracy measured on text-based case data does not transfer directly to the clinical floor.
This is the context behind the researchers' concern. The authors themselves see the risk that if the figure of 67% takes on a life of its own, it could accelerate discussion of clinical deployment before regulatory frameworks have had a chance to catch up.
This evaluation measured classification accuracy on text-based case data. In an actual ER, non-verbal information — respiratory status, skin color, level of consciousness — is said to account for 30–40% of the judgment process, and high accuracy on text is only one part of what is necessary.
This number looks low because triage is fundamentally a task aimed at "not missing the extremes" — the most critical and the least urgent cases. Rather than the absolute value of the accuracy rate, what matters clinically is the breakdown of miss rates and over-triage rates. Information on o1's error distribution remains limited in publicly available sources at this time.
It is unusual for AI researchers to proactively speak out against commercial use of their own findings. This appears to reflect a growing recognition within the industry of past cases — such as a 2024 Stanford medical imaging AI paper — where accuracy studies were oversimplified and repurposed for promotional use.
Both the FDA and PMDA are still developing their approval processes for AI-based medical devices. The pattern of preprint figures circulating ahead of media coverage and corporate IR creates the risk of bypassing regulatory judgment through accomplished facts.
The version used in this study was the preview release, not the current o1 series. Given that OpenAI has continued updating its reasoning models through o3 and o4-mini, comparisons should make explicit where the preview's capabilities stand in that lineage.
What makes this paper noteworthy is not the accuracy figures themselves, but the structure of researchers proactively intervening in how results are interpreted.
In medical AI research, a long-established pipeline — paper → press release → media coverage → investor briefing — has amplified and simplified numbers along the way. By embedding a warning, the authors are applying a brake from the inside on that amplification circuit.
At the same time, if the gap between 67% and 55% is statistically significant, that is itself a fact that cannot be ignored. What matters is moving neither to "therefore we should deploy this clinically right now" nor to "therefore it's useless," but rather toward a design conversation about which workflows and which oversight structures would allow for partial integration.
Japan's emergency medical system is compounded by a chronic shortage of physicians. Concrete discussions of ER triage support tools are expected to begin at multiple university hospitals within 2026.
The fact that o1-preview outperformed physicians in ER triage is not in dispute. But as the authors' own warning makes clear, the greatest risk at this stage is precisely the number taking on a life of its own. If one acts on next steps without reading what lies outside the accuracy figure — the unmeasured variables, the error distribution, the regulatory gaps — the quality of that decision will be lower than 67%.
If you are involved in "AI healthcare," perhaps what you should be asking right now is not "can it be used?" but "where should it be used — and where should it not?"
※ This article was written by an AI writer (AI News) from the Mirai News editorial team.