Google Opens Med-PaLM 2 to Select Customers for Testing

Google planned a limited test of Med-PaLM 2 with select Google Cloud customers to explore possible uses in healthcare. The model scored above 85 percent on U.S. Medical Licensing Examination-style questions, but Google’s team also reported significant gaps in its medical answers.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

The limited healthcare pilot could bring risks from medical errors, but testing and reported limitations keep the story’s overall lean mild.

Google Opens Med-PaLM 2 to Select Customers for Testing

Google is putting its medical language model Med-PaLM 2 into a limited test with select Google Cloud customers. The pilot is intended to explore where the system might be used safely and meaningfully, while researchers continue assessing its strengths and limitations.

What Google says the model can do

Med-PaLM 2 is designed to handle medical questions and information. Google says it could support detailed discussions, respond to complex medical questions, and help find insights in complicated, unstructured medical texts.

The model can produce short or long answers to medical questions. It can also summarize internal documentation, data sets, and scientific sources. These capabilities could help people work with large amounts of medical information, though the announcement describes possibilities rather than established outcomes in clinical practice.

The pilot offers a way for Google and participating customers to explore practical use scenarios. The company says its goal is to examine safe, responsible, and meaningful uses. That emphasis matters because a fluent answer is not, by itself, proof that medical information is complete or reliable.

Exam results show progress, with limits

Google reported that Med-PaLM 2 achieved more than 85 percent accuracy on questions styled after the U.S. Medical Licensing Examination (USMLE). The company described this as expert-level performance and said it was the first language model to reach that level on these questions.

On the MedMCQA dataset, which includes questions from India's AIIMS and NEET medical exams, the model achieved a 72.3 percent pass rate. These scores provide evidence of performance on particular question sets. They do not establish how the system would perform across every kind of medical question or in a real-world care setting.

The earlier Med-PaLM model had reached 67.2 percent on licensing-style questions, where 60 percent was required. Google said the newer model represented an 18 percent increase in performance over its predecessor. Even with that improvement, the team said there was significant room to improve before Med-PaLM 2 meets Google's quality standards.

Evaluation found unresolved gaps

Google assessed Med-PaLM 2 using 14 criteria. These included scientific factuality, accuracy, medical consensus, reasoning, bias, and harm. Clinicians and non-clinicians from diverse backgrounds and countries took part in the evaluation.

The team identified “significant gaps when it comes to answering medical questions,” but the article does not detail what those gaps were. That finding puts the exam results in context: strong performance on a benchmark can coexist with weaknesses that matter when people seek medical information.

The previous Med-PaLM model also illustrates why safety evaluation is part of the picture. Researchers said it produced potentially harmful responses 5.9 percent of the time, compared with 5.7 percent for human experts. Those figures relate to the earlier model and should not be read as results for Med-PaLM 2.

A pilot alongside continued development

Med-PaLM is Google's medical-focused version of PaLM, or Pathways Language Model. The first Med-PaLM was developed using a special soft prompting method combined with responses to medical prompts written by four clinicians. Google later continued development with Med-PaLM 2, but the team did not disclose the technical changes between the versions.

Google plans to work with research teams to address the identified gaps and better understand how the model might improve healthcare. The limited customer test is one part of that exploration. Its stated purpose is to learn about possible applications while the model remains under evaluation.

For now, the reported exam scores show a measurable advance in answering medical questions, while the evaluation findings underline that important limitations remain. The pilot may help clarify how Med-PaLM 2 handles useful tasks with customers, but the source does not report results from that testing.