Artificial intelligence tools are increasingly being explored in mental health care, including as supports for therapy-like conversations, clinician training, and feedback. But in psychotherapy, a response that sounds caring is not always clinically appropriate.
In a recent study, Dr. Venkat Bhat and his colleagues examined whether large language models could generate responses that fit both the conversation and the principles of motivational interviewing, an evidence-based counselling approach that helps people talk through behaviour changes.
Their work points to a central question for the future of mental health care: how can AI-generated therapy responses be tested before they reach patients?
We spoke with Dr. Bhat about what the study found and why human oversight is essential in this process.
How are large language models used in psychotherapy, and how can we test their responses?
VB: Large language models, or LLMs, are being explored as tools that can generate therapy-like conversation, support clinician training, and help evaluate the quality of therapeutic communication.
In this study, we tested this in the context of motivational interviewing. We gave the AI the conversation up to the point where the therapist would speak, then asked it to generate the next therapist response. We compared the AI-generated response with what a human therapist said.
To evaluate the responses, we used two computer-based checks. One measured whether the meaning of the AI response was similar to the human therapist’s response. The other measured whether the response fit the conversation and the therapeutic approach.
What motivated this research?
VB: AI tools are moving quickly into mental health care, but responses that sound supportive are not always clinically appropriate. We wanted to understand whether an AI system could stay aligned with an evidence-based therapy approach, and where it might fall short.
Motivational interviewing was a useful test case because it has clear communication skills, including open questions, affirmations, reflections and summaries. There are also established ways to assess whether therapists are using this approach well.
What was the most important finding of this study, in your opinion?
VB: The most important finding was that the AI often produced responses that fit the conversation but it did not produce the same responses as human therapists.
In other words, an AI-generated response can sound appropriate without being clinically equivalent to a therapist’s response.
The AI performed better when therapist conversations were more focused and consistent. However, its performance declined slightly over longer conversations, and it tended to produce responses that were more verbose than those of human therapists.
How could this change service delivery in the future?
VB: This work supports a cautious path for AI in mental health care.
LLMs may eventually be useful in areas such as therapist training, practice feedback, supervision or clinician-supported digital tools. For example, they may help clinicians review communication patterns or practice therapeutic responses in structured settings.
However, these findings also show that AI should not be introduced as a replacement for therapists. Safe use will require human oversight, clear limits, privacy protections and clinical testing before these tools are used directly in care.
What are the next steps?
VB: The next step is to test these systems using real-world clinical data and human expert reviewers, not only computer-based scores.
We also need to compare different AI models, improve prompts, build therapy-specific evaluation tools and include feedback from patients and clinicians.
A key priority is helping AI stay grounded, personalized, and consistent across longer conversations.
What is the major take-home message for the public?
VB: AI may become a useful partner in mental health care, but “sounds supportive” is not the same as “safe and effective therapy.”
This study shows that AI can generate reasonable therapy-like responses in a structured setting. However, these tools still need careful testing, human oversight and clear boundaries before they are used in patient care.
Special thanks to first author Dr. Bazen Gashaw Teferra and members of the AI for Mental Health team
Read this month's ImPACT paper.
Teferra BG, Huang S, Johny N, et al. Alignment of Large Language Model Responses With Human Therapists in Motivational Interviewing. JAMA Netw Open. 2026;9(3):e262750. doi:10.1001/jamanetworkopen.2026.2750