Scientific papers contain valuable experimental details, but collecting and organizing those details across a large body of research can take substantial time. A UC Berkeley team used ChatGPT to help examine metal-organic frameworks, or MOFs, and reported that the tool completed work that might otherwise have taken a graduate student years.
Turning papers into structured research data
The researchers drew on 228 scientific papers relevant to MOF research. They instructed ChatGPT to extract information, clean it up, and organize it into a structured dataset. The approach focused on gathering details from the literature rather than asking the model to design or conduct experiments.
MOFs are highly porous materials. The source describes their potential relevance to reducing water scarcity and air pollution, making research into these materials important in the context of climate change.
Keeping answers tied to the source
To make the extracted information more reliable, the team designed prompts that directed ChatGPT to use only information provided in each paper. The instructions also told the model to respond with “N/A” when a detail was missing or uncertain.
The prompts narrowed the task further by asking the system to focus on experimental conditions for MOF synthesis and ignore information about synthesizing organic linkers. That scope helped define what counted as relevant information and reduced the chance that unrelated details would enter the dataset.
This source-based approach matters because scientific papers can include many kinds of procedures and results. A system asked to extract specific experimental details needs clear boundaries about which passages to use and what to do when the paper does not provide an answer.
Speed and accuracy in the reported results
The researchers said ChatGPT completed the extraction in “a fraction of an hour,” a task they estimated would have taken a graduate student years. They reported 95% accuracy in extracting data about MOFs.
Those results suggest that language models may help researchers handle literature tasks that are repetitive and time-consuming. In this case, the model was used to turn published information into organized data that chemists could use in further research.
The reported accuracy also makes clear that the system’s output was evaluated against the source material. The workflow was not simply an open-ended conversation: it paired contextual instructions with a defined extraction task.
A wider role for AI in chemistry
The study’s authors concluded that ChatGPT could let chemists process large quantities of information without programming skills. That could shorten the time needed to review research and help scientists move more quickly toward questions about MOFs and their possible applications.
Chayes, dean of the College of Computing, Data Science, and Society, described literature analysis as one important area for “AI for science.” The source quotes Chayes saying the work represents a substantial jump in natural language processing in chemistry, and that a chemist can use the approach without being a computer scientist.
The method may also apply to other areas of chemistry. Its broader lesson is that AI can assist with searching and structuring scientific literature, while the quality of the result depends on defining the task and grounding answers in the papers being examined.