Why medical AI could challenge human-in-the-loop rules

A JAMA opinion piece argues that autonomous AI may soon outperform doctor-AI pairings on medical reasoning tasks. Its authors warn that regulators could lock health care into weaker systems if they require a physician to make the final decision in every workflow.

WTF Index TERMINATOR
◄ Terminator 2 Idiocracy 1 ►

The story mildly leans toward Terminator because it concerns autonomous AI taking over medical reasoning decisions with regulatory and safety implications.

Why medical AI could challenge human-in-the-loop rules

A new argument in JAMA puts a direct challenge to a familiar idea in health care AI: that a doctor should always remain in charge of the final decision. The authors say that rule may sound safer, but could become less safe if autonomous AI systems outperform doctor-AI teams on medical reasoning.

The piece is not arguing that machines are ready to take over every part of medicine. Its focus is narrower: cognitive tasks such as gathering a patient history, diagnosis, test selection, guideline-based treatment, and chronic disease management. Even within that boundary, the authors say regulators should avoid freezing today’s assumptions into tomorrow’s rules.

The Case Against Requiring A Doctor At Every Step

The JAMA piece was led by Ezekiel Emanuel, a bioethicist at the University of Pennsylvania and one of the architects of Obama's healthcare reform. Another author, Neal Khosla, is CEO of the AI telemedicine company Curai Health. His father, Vinod Khosla, is an investor in both OpenAI and Curai Health.

That background matters because the argument points toward a future in which autonomous AI systems have a larger role in care. The source article notes that two of the authors stand to gain directly from the future the piece supports. At the same time, Emanuel’s health policy background gives the piece weight in debates about regulation, liability, payment, and training.

The authors’ position runs against physician groups such as the American Medical Association and the American College of Physicians, which say AI should support doctors rather than replace them. Medical professor Robert Wachter has described AI-only care as the "economy class" of medicine. The JAMA authors reject that ranking as unproven.

What The Evidence Says So Far

The first pillar of the argument is research performance. According to the source, in most studies since 2024, AI alone matches or beats doctors across five core reasoning tasks in medicine: taking a patient history, making diagnoses, choosing tests, treating according to guidelines, and managing chronic disease.

Several examples are central to the claim. Google’s conversational system AMIE scored higher than primary care doctors in almost every category during simulated patient conversations. Across 377 complex cases, ChatGPT o3 named the correct diagnosis first 60 percent of the time, while 20 internists did so 15.9 percent of the time. Microsoft’s diagnostic orchestrator found the correct diagnosis under budget constraints about four times as often as doctors, and at lower cost.

The authors also take a critical view of studies that reach the opposite conclusion. They argue that many are outdated or methodologically weak, including examples that did not use the best available models. Their broader point is that the evidence base is moving quickly, and older comparisons may no longer describe the current systems being debated.

Why Human Oversight Could Become A Weakness

The second pillar is a forecast. The authors expect AI models to keep improving, while doctors may lose some of their own skills as they rely on AI. The source article cites a Lancet study on colonoscopies as a sign of that risk.

If AI becomes clearly better at a task, the authors argue, human review may no longer function as a safety net. It may instead become another place where error enters the process. The issue is not whether doctors are valuable in general, but whether a required human final decision improves the result when the AI system is already stronger at the specific reasoning task.

A meta-analysis of 106 experiments is used to support this point. When the human performs better than the AI, combining them can help. When the AI performs better, the human can make the outcome worse by overruling the system in the wrong places.

One study using real patient cases is especially stark in the source article. GPT-4 alone scored 92 percent on diagnostic reasoning. Doctors using the same model scored 76 percent. For the JAMA authors, that kind of result suggests that mandatory human-in-the-loop rules should not be treated as automatically safer.

The authors also use chess as a comparison. After Deep Blue beat Kasparov in 1997, human-machine teams dominated for years. Starting in 2017, AI began beating those teams too. Their implication is that medicine could pass through a similar stage, where collaboration helps for a while before autonomous systems become stronger than the combination.

The Regulatory Stakes

The policy warning is the heart of the piece. If regulators require a doctor to make the final call in every AI-assisted workflow, they could lock health care into a model that falls behind better-performing autonomous systems. The authors expect that by 2030, autonomous AI will be ready for some, maybe many, workflows, limited to cognitive tasks.

That forecast raises practical questions the piece says should be addressed now. Liability would need to account for decisions made by autonomous systems. Payment models would need to reflect new workflows. Regulation would have to distinguish between cases where human oversight helps and cases where it may reduce accuracy. Medical training would also need to adapt if doctors work alongside increasingly capable systems.

The argument is not that all human involvement should disappear. It is that rules should be flexible enough to follow evidence. A blanket requirement for physician approval may be too blunt if performance differs by task, model, and clinical setting.

Limits The Authors Still Acknowledge

The source article also describes clear limits in the JAMA argument. Almost all the evidence comes from simulations of single tasks, not full real-world patient care. That matters because clinical work involves context, continuity, judgment, and handoffs that may not be captured in controlled comparisons.

The handoff of information between human and model is identified as a weak spot. Even strong reasoning can be undermined if the system receives incomplete, poorly structured, or misunderstood information. The article also separates cognitive work from physical procedures. Surgery, childbirth, and colonoscopies will remain with humans for now because the robotics are not ready.

Autonomous systems also introduce failure modes that doctors do not share. The source names hallucinations, internet outages, and cyberattacks. The authors concede that those risks must be weighed against higher accuracy.

The debate, then, is not a simple contest between doctors and machines. It is about whether health care rules should assume that human review is always the safest default. The JAMA piece argues that this assumption may soon be wrong for some medical reasoning tasks, and that regulators should prepare for that possibility before rules become difficult to change.