I am not especially interested in AI hype, and I’m particularly skeptical of claims regarding AI in healthcare.
I am interested in whether a tool can reduce needless human suffering.
That’s why a new Harvard-backed study in the journal Science related to AI and clinical reasoning hit me a little harder than the average shiny-tech headline.
In a set of experiments comparing a large language model with human physicians, the AI matched or exceeded doctor performance across several reasoning tasks. Most strikingly, in a real-world emergency department comparison using 76 patient cases, the model produced the exact or a very close diagnosis in 67% of triage cases, versus 55% and 50% for two attending physicians reading the same text-based record.
If that finding holds up in broader real-world testing, it matters.
And for me, it feels personal.
My own most recent ER experience is why I cannot shrug this off
My last experience as an ER patient was nightmarish.
I went in with a serious problem: I hadn’t emptied my bladder in over 48 hours. The triage process did not meaningfully capture how bad the situation was, and I was left in escalating pain for over 4 hours.
Eventually, after a long wait and a lot of misery, they hastily pulled me into a side room, made me lay down on a literal bench — not a gurney, not an exam table, apparently all of those were taken: a random little bench barely wide enough to hold me — where they drained roughly 1,400 milliliters of fluid from my bladder.
For reference, the average adult male bladder is considered full at 700 milliliters, and you start getting that “I really have to pee” feeling when your bladder is around half full (i.e. approx. 300-400mls).
That was not a small miss. That was the beginning of a brutal cancer journey that changed my life.

Later, doctors discovered a very large and very rare “embryonal rhabdomyosarcoma”, which is a rare kind of tumor that almost always occurs in children, often in the upper body. I was almost 40, and it was in muscle tissue in my prostate. I did not have prostate cancer, which is all too common. I had a rare pediatric cancer that happened to be in my prostate.
Despite the rocky start, I was fortunate enough to ring the bell in April 2024, and my scans since then have been clear. Although, I’ll still have to go get scanned every 6 months for a while. Regardless, I am deeply grateful to be alive.
That said, I will never forget that ER visit, or the feeling that something had gone badly sideways long before anyone treated it like an emergency.
So, when I read that an AI system may be particularly strong at the exact stage where patients are most vulnerable—early triage, sparse information, high uncertainty—I do not hear, “Cool demo.” I hear a profound opportunity.
I hear, “How many people are suffering right now because the system is overloaded, rushed, inconsistent, or simply too human?”
The real story is not “AI replaces doctors”
Let’s get the obvious part out of the way.
No, this study does not mean machines are ready to replace physicians. The authors themselves were careful about that, and so were outside experts responding to the paper. The study focused on text-based clinical reasoning. It did not test bedside presence, physical exam skill, visual cues, tone of voice, patient distress, or the thousand little judgment calls that happen in a live clinical setting.
That limitation matters a lot when it comes to AI in healthcare.
But it does not make the result trivial. In some ways, it makes it more useful. Plenty of dangerous failures in medicine happen before anyone gets to the elegant part of care. They happen when information is incomplete, the waiting room is full, the chart is messy, people are tired, and a patient is reduced to a few lines of text and a handful of vitals.
That is exactly the kind of environment where the study suggests AI may be able to help.
Why triage is such an important target for AI in healthcare
Emergency medicine is a pressure cooker. People arrive with limited history, incomplete records, vague symptoms, and wildly different levels of urgency. Good triage is difficult. Great triage is harder still.
And the stakes are not theoretical. AHRQ has estimated that, across roughly 130 million annual emergency department visits in the United States, about 7.4 million patients are misdiagnosed, around 2.6 million suffer an adverse event as a result, and about 370,000 experience serious harms tied to diagnostic error.
That does not mean AI is some magical antidote. It does mean even modest improvements in early recognition, escalation, and second-opinion support could matter at enormous scale.
That is the frame I think more people should use.
Not: Can the robot be the doctor?
But: Can the system catch more bad misses before they become disasters?
What this study should wake people up to
The exciting part here is not that an AI model can score well on a medical-school-style quiz. We have had a lot of that already, and most of it is less impressive than the headlines make it sound.
The exciting part is that this research pushed closer to real clinical reasoning.
According to the AAAS summary of the paper, the model consistently matched or exceeded physician performance across six experiments involving diagnosis and management reasoning, with its clearest edge appearing in early emergency department triage under uncertainty. That is a much more interesting result than “AI passed another benchmark.”
Why? Because uncertainty is where patients suffer.
When a person shows up scared, in pain, and not yet fully understood, that is not the moment to rely on a single rushed interpretation of sparse data if a second system could flag danger, suggest alternative diagnoses, or nudge a clinician to look again.
I am not arguing for less human medicine. I am arguing for better-supported human medicine.
Human beings are fallible. That is exactly why tools matter.
One of the stranger habits in public conversation is that we demand perfection from machines while quietly accepting routine failure from human systems.
To be clear, we should demand a lot from AI in healthcare settings. It needs testing. It needs guardrails. It needs auditing, transparency, accountability, and ongoing monitoring. If it performs worse for certain populations, hallucinates, overstates confidence, or introduces new failure modes, that should be exposed quickly and mercilessly.
But we should be just as honest about the status quo. Human clinicians are intelligent, skilled, compassionate, and also tired, overloaded, rushed, and operating inside institutions that are often too big, too complex, and too opaque for their own good.
That is not an insult. That is reality.
And if a tool can help a nurse, resident, attending physician, or hospital team spot the dangerous outlier faster, then refusing to explore it would be its own kind of negligence.
Where AI seems most promising right now
If I were building around this research, I would not start with “fully autonomous diagnosis.” That is the kind of sentence people write right before they wander into a regulatory wall and a moral swamp, especially when it comes to AI in healthcare.
I would start here:
- Triage support: flagging high-risk cases that may deserve faster escalation.
- Second-opinion generation: producing plausible alternate diagnoses from messy chart data.
- Documentation synthesis: pulling key facts out of scattered records so clinicians can think faster.
- Management checklists: surfacing standard next-step considerations without pretending certainty.
- Coverage extension: giving under-resourced settings access to stronger decision support.
Those are not science-fiction use cases. Those are near-term workflow improvements with obvious human value.
The sober takeaway
I can’t say for certain that AI in healthcare would have changed the outcome of my own ER visit, but it could have flagged the severity sooner. Maybe it would have pushed the right data or better questions higher in the queue. Maybe it would have changed nothing at all.
And, either way, the major elephant in the room I haven’t even touched in this article is the fact that medical software is dominated by a grand total of two — although, more practically, only one — company. Getting useful AI in healthcare is, at present, heavily dependent on one of those two giants taking the initiative to do so, which is a story I’ll save for another blog.
Regardless, I do know this: if you have lived through a medical miss, a delay, a shrug, or a long stretch of avoidable suffering, you stop treating “better triage” like an abstract policy goal.
You start treating it like a moral obligation.
That is why this study matters.
Not because AI is ready to replace doctors. It is not.
Not because every hospital should blindly bolt a chatbot onto the emergency department. Absolutely not.
It matters because it offers evidence that AI in healthcare may be genuinely useful — even in one of the hardest, messiest, and most consequential parts of medicine: helping humans make better decisions when time is short and the picture is incomplete.
That is not the end of medicine.
In fact, if we do it right, it could be the beginning of a new golden age of medicine.
If you work in healthcare, here is the question I would ask:
You know that healthcare’s digital transformation has not been all it was promised to be, and you have the annoying EHR and overflowing patient portal inbox to prove it. With that in mind, what digital healthcare nightmares do you most want addressed? If you woke up tomorrow and healthcare’s digital transformation had actually gone as promised, how would your day be different?
Your answers probably hint at where AI in healthcare deserves the most serious look.
If you want a broader picture of how I think about AI adoption, risk, and human opportunity, you might also like Preparing for Careers in the AI Era and my WorkVibe case study on practical, human-centered AI product thinking.
And if you really want to explore how AI in healthcare could improve things for providers and patients alike—or you have an AI-in-healthcare-related idea you want to pressure-test with someone who is enthusiastic about the upside and skeptical of the nonsense—let’s talk.
