AI Therapist Reduced Depression Symptoms by 51%: What the First Clinical Trial Found
In 2025, one of the world's most prestigious medical journals — NEJM AI — published results from the first-ever randomized controlled trial of a generative AI therapist. The headline finding: a 51% reduction in depressive symptoms. This isn't marketing copy — it's peer-reviewed science, and it deserves to be read carefully, both for what it proves and for what it deliberately does not.
The trial in numbers
The study tested a system called Therabot (Heinz et al., 2025, NEJM AI, DOI 10.1056/AIoa2400802). It enrolled 210 adults diagnosed with major depressive disorder, generalized anxiety disorder, or clinically significant eating-disorder symptoms, and randomized them into two arms: four weeks of access to the AI therapist, or a standard waitlist. A four-week follow-up period tracked whether the gains held. On average, participants logged more than six hours of interaction with the system across the trial.
| Arm | N | Condition | Symptom reduction | Effect size (d) |
|---|---|---|---|---|
| Therabot | 106 | Depression | −51% | 0.845–0.903 |
| Therabot | 106 | Anxiety | −31% | 0.794–0.840 |
| Therabot | 106 | Eating-disorder symptoms | −19% | 0.627–0.819 |
| Waitlist | 104 | (no active treatment) | — | — |
Two things make these numbers notable. First, the effect sizes are large — a Cohen's d near 0.85 for depression is in the range clinicians associate with effective in-person therapy, not with a self-help app. Second, participants rated the therapeutic alliance — their sense of being understood and of trusting the interaction — using the WAI-SR scale at around 3.6, comparable to figures typically reported for human therapists. The system was not built casually: it was developed over six years by a team of more than 100 researchers at Dartmouth's Center for Technology and Behavioral Health, using an expert fine-tuned generative model rather than an off-the-shelf chatbot.
The engagement figure is easy to skim past but tells its own story. More than six hours of interaction over four weeks is a serious dose — roughly the equivalent of several therapy sessions' worth of contact — and much of it happened outside conventional clinic hours, when a human therapist simply isn't reachable. The trial also stratified participants across three very different diagnostic groups, and the improvement in eating-disorder symptoms is worth singling out: eating disorders are notoriously hard to shift, and seeing a generative system move that needle at all, even by 19%, was among the more surprising results. None of this settles the debate, but it shows the effect wasn't confined to the single easiest condition.
Why this trial matters
Therabot wasn't the first attempt to build an AI therapist — but it was the first generative one to clear this bar of rigor. The earlier generation was rule-based and scripted. Woebot, launched in 2017, showed a modest but real effect on depression in college students: a Cohen's d of 0.44 in the original trial (Fitzpatrick, Darcy, and Vierhile, 2017), with 83% retention and about 12 sessions on average. It later received an FDA Breakthrough Device designation and entered a larger controlled trial for postpartum depression. Wysa has handled more than 400 million conversations across 65 countries and earned its own FDA Breakthrough designation for chronic-pain-related depression and anxiety.
The controlled evidence for that first wave was real but small. Tess, evaluated by Fulmer and colleagues (2018), produced a significant reduction in depression with a Cohen's d of 0.68 in a trial of 75 people. Youper — an app blending CBT, ACT, and DBT techniques — reported reductions in anxiety (d = 0.60) and depression (d = 0.42) over four weeks in an observational study of more than 4,500 paying users (Mehta et al., 2021). Useful numbers, but none of these systems had been put through a trial with Therabot's methodological weight.
What separates Therabot from all of these is the combination of a generative model — one that composes each response from scratch rather than picking from a scripted decision tree — with a randomized controlled design and effect sizes roughly double those of the rule-based era. A scripted bot can only say what its designers anticipated; a generative one can follow a person into territory no script covered. That flexibility is exactly what makes it more useful and riskier, which is why Therabot's team paired it with constant human monitoring. It is the first serious signal that a system flexible enough to hold an open-ended conversation can also produce clinically meaningful change under controlled conditions — a genuine milestone, and the reason the study drew commentary across the field. Worth noting alongside it: as of late 2025, no generative-AI mental-health tool has received full FDA approval, so even this result lands in a still-forming regulatory landscape.
What it does not prove
The same commentary was quick to flag the study's limits, and taking them seriously is what separates evidence from hype. Three limitations matter most (raised by Bakoyiannis, 2025, Nature Mental Health; Heckman, 2025, NEJM AI; and a 2025 editorial in The Lancet Psychiatry):
- The comparison was a waitlist, not an active treatment. The control group received no care at all. That design can show that Therabot beats nothing, but it cannot show that it matches or beats a human therapist, or even another app. Some of the measured improvement in any waitlist trial reflects attention, expectation, and the passage of time.
- The alliance measure may not transfer. The WAI-SR was designed to capture the bond between a client and a human therapist. Whether a high score against an AI means the same thing — or measures something different, like the comfort of a nonjudgmental interface — is genuinely contested.
- The safety net won't scale as-is. Every conversation in the trial was monitored in real time by the research team. That intensive human oversight is part of why the trial was safe; it is also difficult to reproduce in a product used by thousands of people at once.
None of this erases the result. It frames it: Therabot is strong preliminary evidence, not a finished verdict, and the researchers themselves called for larger trials with active comparators and independent assessment.
How this compares to the meta-analysis
One number puts Therabot in context. Across the broader literature, digital mental-health interventions tend to reach effect sizes of roughly 0.20–0.30, compared with about 0.80 for face-to-face therapy (The Lancet Psychiatry, 2025). Therabot's depression effect sits far above that digital baseline — which is exactly why it drew attention, and also why a single trial can't yet be taken as the new normal.
The most complete picture of the field comes from Linardon and colleagues (2024, World Psychiatry), a meta-analysis of 176 randomized controlled trials of mobile apps for depression and anxiety. Read alongside that body of work, Therabot looks like a possible inflection point rather than an outlier to dismiss — a hint that the generative approach may start closing the gap between digital tools and in-person care. The honest reading is that one strong trial raises the ceiling of what's possible without yet moving the floor of what's proven; replicating it with active comparators is the next real test. We walk through the pooled evidence in our meta-analysis of AI chatbot therapy, and cover the structured-program research it builds on in our review of CBT chatbot studies.
FAQ
Can an AI therapist treat depression?
In one rigorous trial, a generative AI therapist reduced depressive symptoms by 51%, with an effect size comparable to in-person therapy. That is real, peer-reviewed evidence that a well-built system can help. But "help in a monitored four-week study against a waitlist" is not the same as "treat" in the clinical sense — it has not been shown to match a human therapist, and it was run with a level of human oversight most products don't have.
Is AI therapy clinically proven?
Partly, and cautiously. Rule-based tools like Woebot and Wysa have FDA Breakthrough designations and modest trial evidence, and Therabot is the first generative system with a published RCT. But the field's average effect sizes remain below those of face-to-face therapy, and most tools have not been tested at this level of rigor at all. "Promising and improving" is a fairer description than "proven."
Should I replace my therapist with AI?
No. The best current evidence supports AI as a complement, not a replacement — useful for support between sessions, for people on long waitlists, or where no clinician is available. It is not a substitute for a psychologist, psychotherapist, or psychiatrist, and it cannot manage a crisis. If you're in treatment, the sensible move is to use these tools alongside your therapist, not instead of one.
Nearby is a support tool that uses evidence-based psychology. It does not replace a psychologist, psychotherapist, psychiatrist, or emergency service. If you are in crisis, contact your local emergency services or a crisis line right away.