A Skeptical Read of the “AI-Alone Medicine” Argument
https://jamanetwork.com/journals/jama/fullarticle/2852952
There’s a comfortable consensus in medicine about artificial intelligence: the machine assists, the physician decides. The AMA even rebranded it “augmented intelligence” to keep the human firmly in charge.
A recent JAMA Viewpoint by Ezekiel Emanuel and colleagues sets out to demolish that consensus. Its claim is blunt: for the thinking parts of medicine, AI alone will soon be better than a doctor, and better than a doctor working with AI.
It’s a provocative, well-armed argument. It’s also, I think, wrong in some important ways, or at least far less settled than its confident tone suggests.
What the article argues
The authors focus on five “cognitive” medical tasks:
- eliciting a patient’s history
- forming a differential diagnosis
- choosing diagnostic tests
- prescribing guideline-concordant treatment
- managing chronic disease.
Across recent studies, most published since 2024, they marshal evidence that large language models noww match or beat physicians at all five. A few of the headline numbers:
- Google’s AMIE was rated better than primary care physicians at taking a history and at recommending treatment (90% vs 37% favorable on treatment).
- OpenAI’s O3 put the correct diagnosis first in 60% of complex cases; physicians managed under 16%.
- Microsoft’s diagnostic “orchestrator” reached the right answer roughly four times as often as physicians and spent less on testing.
Then comes the genuinely counterintuitive move. The authors argue that once AI alone beats humans alone. Putting a human back in the loop makes things worse, not better.
They cite a 2024 meta-analysis of ~106 human-AI experiments showing that when AI was the stronger solo performer, hybrids underperformed the AI.
The mechanism is “algorithm aversion”, humans distrust and override correct machine judgments, compounded by physician “deskilling” as doctors lean on AI and lose their edge.
Their analogy is chess: centaur teams of human-plus-engine once beat engines alone, until the engines got so strong that human input only added noise. Medicine, they suggest, is on the same trajectory, with AI-alone care ready for real deployment “in some, maybe many, workflows by 2030.”
What it gets right
Give the authors credit for three things.
First, the hybrid-degradation point is real and under-discussed. We reflexively assume “doctor + AI” must beat either alone. That’s an assumption, not a law, and there’s now evidence it fails once the AI is the better solo performer. Anyone designing clinical workflows should take that seriously rather than waving it away.
Second, the piece is honest about its own weak spots. It flags that most studies are simulations, that information transfer between patient and model is a known failure point, and that tail risks (hallucinations, outages, cyberattacks) cut against full autonomy.
Third, the closing call is sound regardless of whether you buy the strong thesis: liability, reimbursement, regulation, and medical education are nowhere near ready for capable autonomous systems, and pretending otherwise helps no one.
Where it wobbles
The evidence is almost entirely from simulations, and the simulations flatter the machine.
Nearly every study feeds the model a clean, text-based vignette in which the relevant facts are already assembled. Real patients don’t arrive as tidy case reports. They arrive with vague complaints, unreliable histories, competing problems, and the crucial detail buried or unmentioned.
The authors themselves cite work showing that patient-to-model information transfer is exactly where LLMs break down. That single caveat undercuts a large share of the accuracy figures, because a model that aces the vignette may fumble the messy encounter that produced it.
The physician baseline is frequently rigged
In the diagnostic-orchestrator study, doctors were barred from consulting colleagues, textbooks, or the internet. That is not medicine; that’s a memory quiz designed to handicap the human.
Many “complex cases” are also drawn from curated teaching collections, the diagnostic puzzles medicine selects because they’re solvable and information-rich. Beating physicians on those is not the same as beating them in a primary care clinic dominated by ambiguity, prevention, and psychosocial context.
“Best care” gets quietly redefined as “task accuracy.”
Reducing medicine to five cognitive tasks is a rhetorical choice that stacks the deck. It brackets out physical examination and procedures (which the authors concede), but also the relational core of care: eliciting what a patient actually values, negotiating uncertainty, building the trust that makes people take the medication at all.
The diabetes result they cite, faster insulin titration with a voice AI, rests on 16 patients per arm. That’s a pilot, not a foundation for “AI is better at chronic disease management.”
The chess analogy is seductive and misleading
Chess is a closed system: perfect information, defined states, an unambiguous win condition. Of course the engine eventually renders the human a liability.
Medicine has none of those properties: noisy inputs, contested ground truth, and consequences for being wrong that don’t exist on a chessboard. The authors admit medicine is “more complex than chess,” then lean on the analogy anyway.
The skepticism is applied asymmetrically
Studies favoring humans are dismissed as outdated or methodologically flawed; studies favoring AI, several produced by the companies selling the systems (Google, Microsoft), are largely taken at face value. That asymmetry deserves a raised eyebrow. And the “algorithm aversion” framing is subtly circular: calling every human override an error assumes the AI was right, which is the very thing in dispute.
Deskilling is treated as evidence for the thesis when it’s really a consequence of it
The argument: AI will pull ahead partly because doctors will deskill. But deskilling is a product of the AI-alone trajectory the authors advocate and it makes their own acknowledged tail risks worse. If physicians atrophy and are then asked to catch a hallucination or cover an outage, the fallback is weaker precisely when you need it most. That’s an argument for keeping humans sharp, not for sidelining them.
The bottom line
The strongest version of this Viewpoint is a useful provocation: stop assuming “doctor plus AI” is automatically the ceiling, and start testing AI-alone against hybrids honestly instead of treating the hybrid as the gold standard by default. That’s a fair challenge to a lazy consensus.
But the leap from “AI wins the vignette” to “AI should run the clinic by 2030” outruns the evidence. The studies measure diagnostic accuracy on curated puzzles under conditions that handicap the human and sanitize the patient, and not outcomes, not trust, not the parts of medicine that don’t fit into five cognitive boxes. The confident timeline is a bet on simulations generalizing to reality, and the history of medical technology is littered with things that worked beautifully in the vignette and disappointed in the exam room.
N.B. This critique was written with the assistance of an LLM
