Artificial intelligence is still not ready to replace scientists. But the work described by researchers at Carnegie Mellon University shows a narrower and more practical possibility: AI systems may be able to take over some of the repetitive, technical work involved in laboratory experimentation.
The system, called “Coscientist,” was designed around cooperation. Instead of asking one model to do everything, the researchers connected three specialized AI instances and gave them separate jobs. Together, they could search for chemical information, understand lab hardware, plan procedures, run experiments, spot errors, and revise their approach.
A divided AI system for chemistry
The researchers wanted to understand what large language models, or LLMs, could contribute to science. The systems used in the work were mostly GPT-3.5 and GPT-4, with Claude 1.3 and Falcon-40B-Instruct also tested. GPT-4 and Claude 1.3 performed the best.
Coscientist was not built as a single all-purpose model. It used a division of labor, with each AI instance responsible for a different part of the workflow. That structure mattered because the task was not just to answer chemistry questions. The system also had to connect that knowledge to real laboratory tools.
The first component was a web searcher. It used Google’s search API to find potentially useful pages, then processed those pages for information. The researchers could see where this module spent its time. About half of the pages it visited were Wikipedia pages, and the top five sites it visited included journals published by the American Chemical Society and the Royal Society of Chemistry.
The second component was a documentation searcher. Its role was to read the manuals for lab automation equipment, including robotic fluid handlers and other devices controlled through specialized commands or a python API. This gave the system a way to learn how the available hardware was supposed to work.
The third component was the planner. It could ask the other two AI instances for help, process their responses, run code in a Python sandbox, and access automated laboratory equipment. In practice, this made the planner the part of Coscientist that acted most like a chemist: gathering information, converting it into a procedure, and trying to carry that procedure out.
From patterns in plates to real reactions
The researchers first tested whether Coscientist could reason about chemical synthesis in general. They asked it to synthesize chemicals such as acetaminophen and ibuprofen. After searching the web and scientific literature, the system could identify viable synthesis routes.
That did not yet prove it could operate a laboratory. To test basic control, the researchers gave it a standard sample plate made up of small wells arranged in a rectangular grid. Coscientist was asked to use colored liquids to fill in squares, diagonal stripes, and other patterns. It handled those tasks effectively.
The next step added a measurement challenge. The researchers placed three different colored solutions at random positions in the grid and asked Coscientist to identify which wells contained which colors. On its own, the system did not know how to solve the problem. After a prompt reminded it that different colors would have different absorption spectra, it used an available spectrograph and identified the colors.
Only after those control tasks did the researchers move to chemistry. They provided a sample plate containing simple chemicals, catalysts, and related materials, then asked the system to perform a specific chemical reaction. Coscientist chose the right chemistry at the start, but its first attempt failed because it sent an invalid command to hardware used for heating and stirring reactions.
That failure became part of the workflow rather than the end of it. The planner returned to the documentation module, corrected the hardware command problem, and ran the reactions. The desired products showed spectral signatures in the reaction mixture, and chromatography confirmed their presence.
Learning from bad guesses
Once Coscientist could run basic reactions, the researchers asked it to improve reaction efficiency. They framed optimization as a game, with the score increasing as the reaction’s yield improved.
The system made poor choices in the first round of test reactions. But after that, it quickly moved toward better yields. The researchers also found that they could reduce those early mistakes by giving Coscientist yield information from a handful of random starting mixtures.
That result matters because it shows how the system could use information from more than one source. Coscientist could learn from reactions it ran itself, but it could also incorporate external information into its planning. In both cases, the information became part of its next experimental decisions.
The researchers identified several major capabilities in the system:
- Planning chemical synthesis using public information
- Finding and processing technical manuals for complex hardware
- Using that knowledge to control different laboratory devices
- Combining hardware control with a laboratory workflow
- Analyzing reactions and using the results to improve reaction conditions
What Coscientist suggests about AI in labs
Coscientist points to a future in which AI helps automate parts of scientific work without replacing human researchers. The system still needed a designed setup, access to raw materials, lab equipment, documentation, and prompts. But once those pieces were available, the researchers could tell it what type of reaction they wanted, and it could work through the steps needed to perform it.
The project also shows why multi-part AI systems may be useful for complex tasks. A single model was not asked to search, read manuals, plan chemistry, write code, control hardware, diagnose software errors, and analyze outcomes all at once. Instead, Coscientist coordinated specialized systems, each handling a clearer slice of the problem.
That structure may help explain why the system could do more than produce text. It was connected to information sources, a Python sandbox, and automated equipment. The result was an AI workflow that could move from research to action, then use the outcome of that action to revise its next plan.
The same capability also raises concern. The researchers were worried about some of what Coscientist could do, because some chemicals, including nerve gasses, should not be made easier to synthesize. The source notes that getting GPT instances to refuse certain tasks remains an ongoing challenge.
For now, Coscientist is best understood as a demonstration of how LLMs can help with scientific labor when they are given the right tools and roles. It does not make AI a scientist. It does show that AI systems can already connect public knowledge, equipment manuals, automated hardware, and experimental feedback in ways that are useful inside a lab.