ChatGPT and other large language models can generate responses that sound human, but researchers say that fluency is not enough to make them suitable psychotherapists. In a paper titled “Using large language models in psychology,” they warn that the systems’ ability to provide psychologically useful information is fundamentally limited.
Human-like language is not human understanding
The researchers’ central concern is that large language models lack a “theory of mind”: an understanding of other people’s mental states. A chatbot may produce language that appears responsive, but that does not mean it understands what a person is feeling or why.
Dora Demszky, a professor of data science in education at the Stanford Graduate School of Education and a co-author of the paper, said the models “are not capable of showing empathy or human understanding.” Her warning draws a line between generating convincing text and providing the kind of understanding that psychological support requires.
David Yeager, a professor of psychology at the University of Texas at Austin and another paper author, makes a similar distinction. A model can sound like a person without having the depth of understanding of a professional psychologist or a good friend.
Why psychotherapy raises particular concerns
Large language models are being considered for uses across fields, including healthcare and psychotherapy. But the researchers argue that psychological applications deserve caution because responses can affect how people think and behave.
That risk is part of what motivates their call for more research. Yeager said the researchers fear “a world in which the makers of generative AI systems are held liable for causing psychological harm because nobody evaluated these systems’ impact on human thinking or behavior.” The concern, as presented in the paper, is that systems could be deployed in consequential settings before their effects have been properly examined.
This does not mean the researchers see no possible role for AI in psychology. Their position is that current capabilities should not be mistaken for readiness to handle the field’s most transformative applications. The gap between potential and demonstrated psychological competence needs to be addressed through research and development.
A shared research effort could set standards
The paper calls for academia and industry to work together on a project on the scale of the Human Genome Project. The proposed collaboration would bring together resources and expertise to develop systems that are more psychologically competent.
Several parts of that effort would help make evaluation more systematic:
- Develop key data sets for research and development.
- Create standardized benchmarks, including benchmarks for psychotherapy use.
- Build shared computing infrastructure to support the development of psychologically competent large language models.
These elements address different needs. Data sets can provide material for studying model performance, while benchmarks offer common ways to assess it. Shared computing infrastructure would support work across the proposed partnership.
Potential depends on further research
The researchers describe both promise and limits. They argue that large language models could advance psychological measurement, experimentation and practice. At the same time, they say the models are not yet ready for many of the most transformative psychological applications.
Their conclusion leaves room for those uses in the future, but ties that possibility to further research and development. For now, the researchers’ warning is that plausible, human-sounding conversation should not be treated as evidence that a chatbot can provide psychotherapy. Establishing what these systems can do safely and competently calls for shared evaluation before wider use in psychological practice.