How BioGPT focuses language modeling on biomedical research

Microsoft’s BioGPT was trained on biomedical article titles and abstracts to handle tasks such as question answering and relation extraction. The research team reported strong results on biomedical benchmarks, including for a larger version of the model.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 1 ►

This is a routine report on a specialized model and benchmark results, with no clear lean toward either outcome.

How BioGPT focuses language modeling on biomedical research

Microsoft developed BioGPT as a language model focused on biomedical work. Rather than training it on broad general text, the research team used biomedical publications and evaluated the model on tasks including question answering, text generation, relation extraction, and document classification.

Training on biomedical publications

The team gathered titles and abstracts from PubMed, an English-language text-based meta-database of biomedical articles. The collection was updated before 2021 and contained 15 million pieces of content for training.

BioGPT is based on GPT-2 and has 357 million parameters. The team pre-trained it using eight Nvidia V100 GPUs for 200,000 steps, then fine-tuned it with a single Nvidia V100 GPU for 32 steps. Fine-tuning adapted the model for particular biomedical tasks after its initial training.

That domain-specific focus shaped what BioGPT was designed to do. Its downstream tasks included end-to-end relation extraction, text generation, question answering, and document classification. Together, these cover both producing biomedical language and working with information in biomedical texts.

How it compared on benchmarks

According to the research team, BioGPT outperformed comparable models based on Google BERT on biomedical question answering and end-to-end relation extraction benchmarks. In biomedical text generation, the team also reported better results than a generally trained GPT-2.

The article illustrates the difference with a prompt about the treatment of COVID-19. GPT-2 produced an answer that mixed in unrelated claims and references to COVID-20 and COVID-22. BioGPT generated a response about remdesivir and SARS-CoV-2, though the example alone does not establish how reliable either model would be across other prompts.

The researchers also scaled their GPT-2 medium-based model to the largest GPT-2 XL architecture available. Their fine-tuned BioGPTLarge had 1.5 billion parameters and achieved 81 percent accuracy in the PubMedQA benchmark, compared with 78.2 for BioGPT.

In that benchmark, BioGPTLarge also scored above Flan-PaLM, which had 540 billion parameters and a score of 79.0, and Metas Galactica, which had 120 billion parameters and a score of 77.6. These results suggest that a model trained for a narrower subject can compete with much larger general models on tasks within that subject area.

What the results suggest

The findings point to two ways to adapt language models for specialist work. One is to train a comparatively small model on domain-specific material, as Microsoft did with BioGPT. The article also describes fine-tuning large models such as PaLM for specialist use; Google’s Med-PaLM was presented as an example, using specialized prompts and high-quality data to answer lay medical questions at the level of human experts.

Microsoft Research said BioGPT performed at the level of human experts on the tasks tested in the benchmarks and outperformed other general and scientific language models. The team said the model could help researchers gain insights in areas such as drug development or clinical therapies. Those claims relate to the evaluated tasks and do not, by themselves, show that generated text should be treated as medical guidance.

Further work planned

The team said it planned to test further scaling, with a larger BioGPT trained on more biomedical data and optimized for more tasks. The direction reflects an ongoing trade-off: domain-specific training can focus a model on a field, while broader systems can be adapted with specialized data and prompts.

BioGPT’s benchmark results make the case for measuring a model against the work it is meant to perform, rather than parameter count alone. For biomedical research, question answering, relation extraction, and text generation each offer different ways to assess whether a language model can use specialist material effectively.