Google and DeepMind researchers have adapted a large language model to answer medical questions from the public. Their Med-PaLM system performed much better than its untuned counterpart in several evaluations, with some results close to clinician answers. The work suggests a possible role for medical AI, while also showing why strong test results do not make a model a substitute for medical professionals.
Adapting a general model for medical questions
Med-PaLM builds on PaLM, Google's language model with 540 billion parameters. The research team started with Flan-PaLM, a PaLM variant trained to follow instructions across tasks such as dialogue, frequently asked questions, and reasoning.
Rather than retraining the model on a large collection of medical data, the researchers developed a method called Instruction Prompt Tuning. It combines learned soft prompts with human-written prompts designed to guide responses to medical questions. Four clinicians from the US and UK helped create the human-written prompts.
The researchers describe the approach as efficient in its use of data and model parameters. In practical terms, the method aims to steer a broad language model toward medical responses without carrying out the more complex process of fine-tuning it directly with medical data.
How Med-PaLM performed
Clinicians assessed the quality of the model's answers. Med-PaLM substantially outperformed Flan-PaLM on medical responses, and on several measures its performance came close to that of professionals. The results were not uniformly equivalent to clinician performance, however, and the research team said the findings were encouraging while still falling short of clinicians overall.
One important measure was whether responses aligned with scientific consensus. According to the researchers, 92.6% of Med-PaLM answers were judged to meet that standard. Clinician responses scored 92.9%, while Flan-PaLM reached 61.9%. Those figures point to a substantial improvement after the medical adaptation.
The evaluation also considered answers that could potentially cause harm. Such responses accounted for 29.7 percent of Flan-PaLM's answers, compared with 5.9 percent for Med-PaLM and 5.7 percent for human experts. The gap between the adapted model and the unadjusted version suggests that how a model is guided can matter alongside its underlying scale.
Useful progress, with limits
When laypeople compared answers, they rated human experts as more helpful than either language model. Med-PaLM nevertheless did significantly better than Flan-PaLM. This distinction matters: an answer can be closer to scientific consensus or less likely to be harmful and still not feel as helpful to the person asking the question.
The researchers also found that performance improved across PaLM models ranging from eight to 540 billion parameters. They suggest that Med-PaLM's capabilities may emerge as models grow. But their findings also show that scale by itself is not enough: Flan-PaLM performed comparatively weakly, while Instruction Prompt Tuning was used to improve its medical responses.
Together, the results indicate that careful prompting may help improve accuracy, factual consistency, safety, and other qualities relevant to medical AI. They do not establish that a model can reliably handle every real-world health question. The researchers frame the work as a step toward clinical applications and as a reason for further discussion about developing medical AI that is easier, safer, and more equitable to use.
A benchmark for medical question answering
Alongside Med-PaLM, the team introduced MultiMedQA, a benchmark combining six existing open-ended datasets. It covers questions related to medical exams, research, and consumer inquiries. The researchers also created HealthSearchQA, a free-text dataset of medical questions searched online.
These resources give researchers ways to evaluate systems across different kinds of medical questions. The findings offer evidence that adapting a general-purpose model can make its answers more scientifically aligned and reduce potentially harmful responses, while the human-expert ratings underline that performance still has limits.