ChatGPT’s performance on complex clinical questions is pushing medical educators to reconsider how they teach and assess future doctors. A Stanford study found that the AI outscored first- and second-year medical students, on average, on the case report portion of an exam.
Testing clinical reasoning beyond multiple choice
Earlier studies had examined ChatGPT’s ability to answer multiple-choice questions on the United States Medical License Examination (USMLE). The Stanford researchers focused instead on open-ended questions, which are used to assess clinical reasoning skills.
On average, the AI model scored more than four points higher than medical students on the case report portion. The study, published in JAMA Internal Medicine, also reported a significant improvement from GPT-3.5, which was described as “borderline passing” on the questions.
That performance matters because written case questions ask students to explain how they approach a clinical scenario, rather than select an answer from a list. If an AI system can perform well on that kind of assessment, educators may need to reconsider what written exams measure and how students demonstrate their reasoning.
Teaching doctors to think and use AI
The study does not remove the need for doctors to reason through cases themselves. Alicia DiGiammarino, education manager at the School of Medicine and a co-author, described both sides of the challenge: students should not become so dependent on AI that they fail to develop independent reasoning, but they also need training to use tools that are appearing in modern practice.
This creates a curriculum question for medical schools. Students need practice making judgments without relying on an AI response, while also learning how to incorporate AI into their work. Those aims can sit together if education treats the technology as a tool to assess and question, rather than as a substitute for clinical reasoning.
AI performance on an exam is also not the same as safe clinical practice. The article identifies invented facts, known as hallucinations or confabulations, as a major weakness. Even occasional errors can have serious consequences in medical contexts.
The article says this problem has been significantly reduced in OpenAI’s latest model, GPT-4, but remains present. It suggests that using AI within a curriculum that draws on multiple sources of truth could make the risk smaller. That still requires students and doctors to check what a system says rather than accept it uncritically.
Stanford adjusts exams while exploring AI
At Stanford’s School of Medicine, concerns about exam integrity and AI’s effect on curriculum design have already led to changes. Administrators switched from open-book to closed-book exams so students would develop clinical reasoning skills without relying on AI during the assessment.
At the same time, the school created an AI working group to explore how these tools might be integrated into medical education. The two steps point to a practical tension: students need settings where they can demonstrate their own reasoning, and educators need to decide how AI belongs in learning and practice.
That balance will shape what medical assessments are designed to test. A closed-book exam can show what students can do without outside assistance, while a curriculum that addresses AI can prepare them to use it with care. The Stanford findings make that distinction harder to ignore.
AI’s reach extends beyond the classroom
The article describes other developments in healthcare that put medical education in a broader context. Insilico Medicine recently administered the first dose of a generative AI drug to patients in a Phase II clinical trial. Google is field-testing Med-PaLM 2, a version of its large language model PaLM 2 fine-tuned to answer medical questions.
It also points to a study suggesting GPT-4 can help doctors answer patients’ questions with more detail and empathy. These examples do not establish that AI can replace a doctor’s judgment. They show why future physicians may encounter AI across more than one part of healthcare, and why training must address both its capabilities and its limits.
For medical schools, the task is not simply to decide whether AI belongs in education. The Stanford study raises a more specific question: how can students learn to use AI while still building the clinical reasoning that their work depends on? Preparing them for both demands may require changes to teaching and testing alike.