Google is putting Med-PaLM 2, its medical language model, into early clinical trials with healthcare customers. The tests bring a system built to answer medical questions into settings where its capabilities and limitations matter directly to patient care.
What the model is being tested to do
Med-PaLM 2 is a medical version of Google’s PaLM language model. It was trained on questions and answers from medical licensing exams to strengthen its ability to respond to medical questions.
The system can also summarize medical documents, organize health data and generate answers to medical questions. Those tasks point to possible support for handling health information, though they do not establish that the model can replace a clinician’s judgment.
According to the Wall Street Journal, initial testing is underway at healthcare facilities in the U.S., including the Mayo Clinic. Google has said the technology could be especially useful in countries with “limited access to doctors.”
Exam results show promise, with limits
Google announced the first clinical trials of Med-PaLM 2 in April of this year. The company says the model performs 18 percent better than its predecessor and outperforms similar models on medical tasks.
Google also says Med-PaLM 2 is the first language model to exceed 85 percent accuracy on questions similar to those on the U.S. Medical Licensing Examination (USMLE). On the MedMCQA dataset, which includes questions from India’s AIIMS and NEET medical entrance exams, it earned a “satisfactory score” of 72.3 percent.
These results offer measures of performance on exam-style questions. They do not remove the possibility of errors when the model is used to produce medical information in practice.
Clinical use raises questions about trust
Google researcher Greg Corrado, who helped develop Med-PaLM 2, has described it as a technology he would not yet use for his family’s health care. The caution sits alongside his view that the model expands the possibilities of AI in medicine tenfold.
Together, those points capture the tension in the trial: a model may be capable of useful medical tasks while still falling short of the confidence required for personal health decisions. Testing in healthcare facilities can explore where it helps, while the acknowledged risk of mistakes remains relevant.
Patient data and AI advice remain concerns
Handling sensitive patient information is one concern as AI moves into healthcare. The Wall Street Journal reports that data submitted by customers during the Med-PaLM 2 trial will be encrypted, inaccessible to Google and controlled by the customers themselves.
Safeguards for submitted data address one part of the discussion. The potential risks of AI-generated medical advice are another, especially when a system can produce answers that sound authoritative but may still be wrong.
Google introduced the first Med-PaLM at the end of 2022. A study published at the end of April 2023 also found that even a non-medically fine-tuned version of ChatGPT based on GPT 3.5 could receive higher ratings for quality and empathy than physician responses when evaluated by humans. That result adds context to interest in AI responses, but it does not settle how such systems should be used in care.
Med-PaLM 2’s clinical testing therefore combines practical potential with unresolved limits: it can process medical information and has reported strong exam results, yet mistakes, patient privacy and the role of human judgment remain central considerations.