Can AI Be a Therapist? What a Stanford Study Found
A Stanford study found that psychiatrists often disagree about whether AI mental health responses are safe. Here is what that means for AI therapy.
People increasingly talk to AI about anxiety, loneliness, relationships, trauma, and depression.
The appeal is obvious. An AI chatbot is available at any hour. It does not appear impatient or judgmental. It can ask questions, remember details, and respond in language that feels attentive and empathetic.
For many users, the experience can feel surprisingly close to therapy.
But sounding therapeutic and being a therapist are not the same thing.
A recent Stanford study examined how psychiatrists evaluate AI-generated mental health responses. The researchers found that experts often disagreed about whether the same answer was safe, empathetic, correct, or appropriate.
The disagreement was especially pronounced in situations involving suicide and self-harm, where mistakes may have the most serious consequences.
The problem was not simply that the psychiatrists were inconsistent. They were applying different but coherent ideas about what a good mental health response should accomplish.
Before developers can train an AI to respond correctly, they must first decide what “correctly” means.
What the Stanford researchers tested
The study, Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing, was accepted at the 2026 ACM Conference on Fairness, Accountability, and Transparency.
The researchers asked three board-certified psychiatrists to independently assess 360 responses generated by large language models. The responses addressed synthetic mental health prompts rather than real patient data.
The prompts covered topics including:
- depression
- anxiety
- psychosis
- addiction
- eating disorders
- obsessive-compulsive disorder
- ADHD
- suicidal ideation
- self-harm
Each response was evaluated across eight dimensions, including correctness, harm, empathy, relevance, professional boundaries, and actionability.
Together, the psychiatrists produced 1,080 evaluations. They received the same rubric and completed a calibration process before beginning.
Despite that preparation, their agreement remained consistently poor.
The study reported inter-rater reliability scores between 0.087 and 0.295, below the 0.40 threshold the researchers considered minimally acceptable for consequential assessments.
This matters because expert ratings are often treated as the ground truth used to train, compare, and approve AI systems.
When experts cannot reliably agree, it becomes unclear what the model is supposed to learn.
Why the psychiatrists disagreed
After the evaluation, the researchers interviewed the psychiatrists about how they made their decisions.
They identified three broad clinical orientations.
A safety-first approach
One psychiatrist focused primarily on whether a response could directly cause harm.
From this perspective, an answer could be considered relatively successful if it avoided dangerous advice, inappropriate certainty, or encouragement of harmful behavior.
Emotional warmth was valuable, but secondary to immediate safety.
An engagement-centered approach
Another psychiatrist placed more importance on empathy, practical support, and keeping the user engaged.
A technically cautious response could still be judged poorly if it felt generic, cold, or dismissive.
Repeatedly telling a vulnerable person to seek professional help may reduce immediate risk, but it may also make them stop talking.
A culturally informed approach
The third psychiatrist focused more strongly on social context, cultural assumptions, boundaries, and access to care.
Advice that sounds appropriate for one person may be unrealistic or alienating for another. It may assume access to therapy, financial stability, a supportive family, or a particular cultural understanding of mental health.
None of these approaches is obviously irrational.
Each protects against a different kind of failure, and each may recommend a different response to the same situation.
Why averaging expert ratings is not enough
A common response to disagreement is to collect more ratings and calculate an average.
Mathematically, this produces a single score.
Clinically, it may represent no one’s actual judgment.
One psychiatrist may believe that caution is necessary because the user could be in danger. Another may believe that the same caution is harmful because it breaks trust and discourages further disclosure.
Averaging their scores does not resolve that conflict. It removes the reasoning behind it.
The resulting AI may learn an unstable compromise between incompatible goals.
It may sound empathetic without becoming meaningfully engaged. It may recommend professional help without explaining why. It may avoid an obviously harmful statement while still failing to respond appropriately to the user’s distress.
The Stanford researchers describe these aggregated judgments as arithmetic compromises. Instead of representing a coherent clinical philosophy, the average may erase the philosophies behind the individual ratings.
The question is therefore not only how many experts should evaluate an AI. It is also which priorities and values the system should represent.
“Safe” is not a single quality
AI safety in mental health is often treated as if it were one measurable property.
In reality, several different risks may exist at once.
A chatbot can cause harm by:
- reinforcing a delusion
- validating an inaccurate interpretation
- failing to recognize an emergency
- responding so cautiously that the user disengages
- offering advice without enough context
- creating false confidence in a diagnosis
- encouraging emotional dependence
- making cultural or social assumptions
- delaying contact with qualified care
Reducing one risk can increase another.
A chatbot that refuses to discuss anything related to self-harm may avoid saying something dangerous. It may also become unusable for someone who genuinely needs to discuss distressing thoughts.
A chatbot that responds with warmth and validation may keep someone engaged. It may also confirm paranoia, resentment, or an inaccurate interpretation of events.
A chatbot that repeatedly recommends professional care may be acting cautiously. But that advice may be practically useless to someone who cannot afford or access treatment.
There is no universal response strategy that maximizes every form of safety at once.
Empathy can also mislead
One of the most persuasive qualities of conversational AI is its ability to sound empathetic.
But generated empathy can make an uncertain answer feel authoritative.
Imagine someone who believes that a friend deliberately ignored them. An AI might respond:
It makes sense that you feel betrayed. They clearly did not respect your feelings.
The response sounds supportive, but the chatbot does not know what the friend intended. It has validated both the emotion and an unverified interpretation.
A more careful answer might say:
It makes sense that you felt hurt. From what you described, though, their intention is not completely clear. What other explanations might fit what happened?
The second answer preserves uncertainty.
This distinction is central to mental health support. A person’s emotions can be real and understandable even when their interpretation of a situation is incomplete or mistaken.
An AI optimized mainly for approval may struggle with this boundary. Agreement often feels more supportive than disagreement, and confident interpretations often feel more insightful than cautious ones.
A chatbot can therefore appear helpful while reinforcing beliefs that should instead be examined.
Therapy is more than a good conversation
The Stanford study evaluates individual responses, but therapy is not a collection of isolated answers.
It is a process that develops over time.
A therapist forms hypotheses, notices patterns, asks for missing information, remembers earlier sessions, and adapts to changes in the patient’s condition.
The immediate emotional effect of a response is not always the same as its long-term therapeutic value.
Reassurance may reduce distress while reinforcing reassurance-seeking. Challenging a belief may initially feel uncomfortable while eventually helping the person examine it more clearly. Recommending urgent care may be necessary even if the patient experiences it as a rupture.
An evaluator looking at one chatbot response cannot observe those longer-term effects.
A general-purpose AI chatbot faces the same limitation. It may remember previous messages, but memory is not the same as clinical understanding. It cannot observe the user’s physical condition, environment, relationships, or behavior outside the conversation.
It also does not carry responsibility for the outcome in the way a licensed professional does.
Calling an AI a therapist risks reducing therapy to its most visible feature: language that sounds psychologically informed.
The issue extends beyond therapy apps
Many people do not use products explicitly marketed as AI therapists.
They use general-purpose assistants, companion apps, and role-playing chatbots for emotional support.
A company may state that its product is not intended for mental health care, while users still bring it highly personal problems. The chatbot can become part adviser, part confidant, and part emotional companion without ever formally entering a clinical setting.
That makes oversight difficult.
Any conversational system capable of sustaining intimate and personal interactions may eventually be used for mental health support, whether or not its developers intended that use.
Safety cannot therefore be limited to products that call themselves therapy apps.
What the study does not prove
The Stanford study does not show that AI can never provide useful mental health support.
It also does not compare AI conversations with professional therapy, peer support, self-help resources, or receiving no support at all.
The study involved only three psychiatrists, synthetic prompts, and isolated model responses. The researchers describe their work as an existence proof rather than a definitive measurement of agreement across the mental health profession.
Different experts or evaluation systems might produce different results.
The study also examines expert judgment rather than long-term patient outcomes. It shows that psychiatrists disagree about responses, but not how users would ultimately be affected by those responses.
Its conclusion is narrower:
Expert evaluation does not automatically produce a single, objective definition of safe AI behavior in mental health.
That is still important because much of AI training and safety testing assumes that expert ratings can be combined into stable ground truth.
So, can AI be a therapist?
A general-purpose chatbot should not be treated as a therapist.
It may reproduce some therapeutic techniques. It can ask reflective questions, explain concepts, suggest exercises, and help someone describe what they are experiencing.
But producing therapeutic language is not the same as delivering therapy.
An AI chatbot cannot reliably diagnose a condition, assess every relevant risk, understand the user’s full context, or take responsibility for the consequences of its advice.
That does not make AI useless in mental health.
AI may support professionals by summarizing information, assisting with administrative work, helping deliver structured interventions, or supporting training. It may also help users prepare for appointments or better understand general psychological concepts.
The important distinction is between supporting mental health care and replacing the professional responsible for providing it.
The first is plausible.
The second is not supported by the capabilities and safety evidence available today.
What responsible systems would require
The Stanford researchers argue that developers should preserve expert disagreement rather than averaging it away.
A model could be evaluated separately under different frameworks. A response might perform well under a safety-first approach but poorly under an engagement-centered or culturally informed one.
That would be more informative than reducing everything to one score.
Responsible AI mental health systems would also require:
- transparent evaluation methods
- testing across cultures and populations
- participation from people with lived experience
- clear crisis escalation procedures
- long-term monitoring
- honest communication about limitations
- accountability when failures occur
- evidence based on real outcomes, not only expert impressions
No benchmark can remove every disagreement from mental health care.
The goal should not be to pretend that one universal answer exists. It should be to make the assumptions behind the system visible and limit its authority where evidence remains uncertain.
AI can sound therapeutic before it is safe enough to be therapy
The main lesson from the Stanford study is not simply that experts disagree.
It is that AI mental health safety cannot be reduced to collecting enough ratings and averaging the answers.
The experts may be protecting against different risks. Their disagreement may reveal genuine uncertainty about what the system should do.
AI chatbots can already sound calm, empathetic, insightful, and reassuring. Those qualities make them useful, but they also make them easy to overestimate.
The ability to imitate the language of therapy arrived before the evidence, accountability, and clinical understanding required to replace a therapist.
Until that gap closes, AI should be treated as a tool that may support parts of mental health care, not as an autonomous therapist.
Important: AI chatbots are not substitutes for diagnosis, treatment, or emergency care. Anyone who may be in immediate danger should contact local emergency or crisis services rather than relying on an AI conversation.
Sources
- Expert Evaluation and the Limits of Human Feedback in Mental Health AI Safety Testing, Kiana Jafari et al., accepted at ACM FAccT 2026.
- Stanford Study Exposes Major Flaw in AI Mental Health Safety Testing, Stanford Institute for Human-Centered Artificial Intelligence, July 13, 2026.
- Towards responsible AI for mental health and well-being: experts chart a way forward, World Health Organization, March 20, 2026.