AI in mental health: what the evidence shows and where the boundaries are

On artificial intelligence and mental health there are two equally simplified accounts circulating. One says that AI therapy is just around the corner and will replace psychologists. The other says that it is all dangerous smoke and that there is nothing to discuss.
The research of recent years does not support either. There is evidence of efficacy in confined contexts and, at the same time, solid evidence of serious failures in critical clinical situations. It is worth looking at the studies in particular, because the practical conclusion for those exercising is quite clear and is not "all or nothing".
TL;DR
- The first randomized trial of a generational AI chatbot for mental health treatment (Therabot, 210 participants) showed symptomatic reductions from the waiting list: 51% in depression, 31% in generalized anxiety, 19% in body image concerns.
- The authors themselves conclude that no generative AI agent is ready to operate autonomously in mental health, and the trial had human clinical supervision.
- A meta-analysis of conversational agents (35 studies, 15 randomized trials) found moderate effects on depressive symptoms (g = 0.64) and discomfort (g = 0,70).
- In contrast, a 2025 study showed that language models express stigma towards pictures such as schizophrenia and alcohol dependence, and failed to recognize crises, even encouraging delusional ideation.
- The place where evidence and common sense coincide today: AI as a professional support tool, not as a substitute for treatment.
Why it Matters
Your patients are already using AI. They tell things to a chatbot at 3:00 in the morning, ask for interpretations of dreams, consult if what they feel is "normal" before or after talking to you. It's not a future possibility, it's current behavior.
Knowing what the research shows allows you two things: to respond with judgment when a patient asks you, and to decide with your own judgment what tools you will incorporate into your practice and which ones you will not.
What the evidence says
The First Randomized Trial
In March of 2025, Heinz et al. published in NEJM AI the first randomized controlled trial of a generational AI chatbot designed for mental health treatment (Therabot, developed in Dartmouth).
The design: 210 adults with clinically significant symptoms of major depression, generalized anxiety or high risk of eating disorder. 106 received access to the app for four weeks, with four additional weeks of optional use; 104 remained on the waiting list.
Results vs. control:
- Depression: average reduction in symptoms of 51%.
- Generalised anxiety: 31%.
- Weight and body image concerns: 19%.
Participants used the app for a total of about six hours throughout the trial, approximately the equivalent of eight sessions.
Important
The comparator was waiting list, not therapy with a professional. That means the study shows that the chatbot is better than receiving nothing, not comparable to psychological treatment. It is a central distinction and is usually lost in headlines.
In addition, the trial was conducted with clinical supervision: there were people reviewing the interactions and ability to intervene. The researchers themselves stressed that no generative AI agent is currently in a position to operate completely autonomously in mental health.
The accumulated evidence of chatbots
Prior to Therabot there was already a research body on conversational agents, mostly based on rules or scripts (not generative AI). The meta-analysis of Li et al. (2023), in npj Digital Medicine, reviewed 35 studies and included 15 randomized trials.
It found significant reductions in:
- Depressive symptoms: g = 0.64 (IC 95% [0.17; 1.12]).
- Psychological distress: g = 0.70 (IC 95% [0.18; 1.22]).
They are moderate effects, with broad confidence intervals, in typically short interventions and low access threshold.
The other side: failures in critical situations
Moore et al. (2025) presented a study at the ACM FAccT conference that evaluated five language models and five commercial "therapy" bots against clinical scenarios.
The findings are worrying:
- Models expressed more stigma toward alcohol dependence and schizophrenia than to depression, and that stigma did not improve in newer or larger models.
- Faced with crisis scenarios, the models failed to recognize the risk.
- In several cases they accompanied delusional thinking rather than contrasting it with reality, which contradicts established clinical practice.
The authors' conclusion is explicit: these flaws prevent language models from safely replacing mental health professionals.
What does it mean in your clinical practice
- I distinguished two completely different uses. One thing is patient-oriented AI as a substitute for treatment (where evidence is incipient and risks are documented) and another is AI as a professional support tool: transcription, draft notes, organization of clinical information. They are separate debates and should not be mixed.
- Ask actively for the use of AI. If a patient consults a chatbot between sessions, that is clinical material: what he seeks there, what he finds, what he cannot tell you.
- Be clear about crisis limits. If you work with patients at risk, it is worth explaining that a chatbot is not an emergency resource and writing down which ones are.
- Apply the same confidentiality criteria as any other tool. To dump identifiable clinical material into a general-purpose chatbot involves sharing health data with a third party, with specific legal consequences. We develop it in ChatGPT and clinical data.
- No AI replaces your judgment. A draft note is a draft: you review it, correct it and sign it you, who was in the session.
Council
When a patient says "I asked the AI," there's usually something more interesting behind it: why it was easier to ask a machine. That's the conversation it's worth, much more than evaluating whether the answer it received was correct.
Killings and limitations
- A trial does not make a body of evidence. The study by Therabot is a methodological milestone, but it is only one, brief, with waiting list control and in population that for the most part did not receive another treatment.
- Follow-up is short. Four to eight weeks say nothing about sustaining medium-term results.
- The systems evaluated are not the commercial systems. Therabot was developed and supervised by an academic team; most available apps have neither that design nor that supervision.
- Technology changes faster than research. Any findings about specific models age in months, although Moore et al.'s study suggests that certain failures are not solved alone with newer models.
- There is almost no evidence in Spanish or Latin American population.
In summary
The evidence available draws a reasonably sharp line. There are signs that well-designed and supervised conversational tools can help people who today do not have access to anything. And there is hard evidence that current models fail just where the most expensive fail: crisis, risk, severe pictures.
Therefore our position on clinical AI is the same from day one: the therapeutic work is yours and is not delegated. What can be delegated is the administrative burden that steals time and attention. Brauni prepares the draft of the note from the session, with your data accommodated safely and not used to train models, so that you decide what is registered.
References
- Heinz, M. V., Mackin, D. M., Trudeau, B. M., Bhattacharya, S., Wang, Y., Banta, H. A., ... Jacobson, N. C. (2025). Randomized trial of a generative AI chatbot for mental health treatment. NEJM AI, 2(4). doi.org/10.1056/AIoa2400802
- Li, H., Zhang, R., Lee, Y. C., Kraut, R. E., & Mohr, D. C. (2023). Systematic review and meta-analysis of AI-based conversational agents for promoting mental health and well-being. npj Digital Medicine, 6, 236. doi.org/10.1038/s41746-023-00979-5
- Moore, J., Grabb, D., Agnew, W., Klyman, K., Chancellor, S., Ong, D. C., & Haber, N. (2025). Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers. Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT '25). doi.org/10.1145/3715275.3732039
Free Brauni test for 30 days, no card
Automatic session notes, digital medical records and more.
Start for freeRelated articles

Based on Evidence
How many sessions do you need? Evidence of dose and frequency in therapy
Dose-response research shows improvement curves, optimal session ranges, and an uncomfortable finding of weekly frequency. What does it mean when planning treatments.

Based on Evidence
Treatment of Trauma: What the Evidence Says About EMDR and Trauma-Focused TCC
A network of 90 trials and 6.560 patients compared 22 treatments for post-traumatic stress. What EMDR and trauma-centered CCT show, and what it means when choosing approach.

Based on Evidence
Does online therapy work just like face-to-face therapy? What does the evidence say
Meta-analysis with thousands of patients compare video and face-to-face psychotherapy: equivalent results and therapeutic alliance without differences. What does it imply for your clinical practice.