Skip to main content
AI & modern therapy

ChatGPT as Therapist: Opportunities and Risks of Large Language Models

By Nearby Published on March 2, 2026 Updated on July 5, 2026 8 min read

Millions of people already discuss their anxieties, fears, and personal struggles with ChatGPT and other language models. This happens spontaneously, without clinical protocols or oversight. It's worth asking honestly what a general-purpose chatbot does well in that role, where it fails, and why the difference matters.

If you are in crisis or thinking about harming yourself, a general chatbot is not the place to turn. Contact your local emergency services or a crisis line immediately — real, trained human help is available right now.

Why people turn ChatGPT into a therapist

The appeal is easy to understand. ChatGPT is available at 3 a.m. on a bad night, it costs nothing or almost nothing, there is no waitlist, and there is no receptionist, no referral, and no fear of being judged. For someone who has never spoken to a therapist — or who can't afford or reach one — typing a worry into a chat box is the lowest-friction first step imaginable. That accessibility is real and shouldn't be dismissed. But "easy to reach" and "safe to rely on" are different claims, and most of the value and most of the danger come from the same underlying trait: a general LLM will say something fluent about anything you ask.

The scale of this is worth pausing on. People are turning to general chatbots for emotional support in enormous numbers, entirely outside any clinical system — a market for digital mental-health tools that reached roughly $11 billion in 2025. That volume is a signal of genuine unmet need: therapy is expensive, waitlists are long, and stigma is real. It is also a signal of risk, because the millions of these conversations happen with no protocol, no oversight, and no one checking whether the advice given was sound. The right question is not "should people do this?" — they already are — but "what has to be true for it to be safe?"

What general LLMs do reasonably well

Within limits, large language models — systems like ChatGPT, trained on vast amounts of text — genuinely help. They can sustain a supportive dialogue, reflect back what a person said so they feel heard, explain cognitive-behavioral techniques in plain language, and walk someone through a structured exercise like reframing a catastrophic thought or breaking a worry into a thought record. Ask one to explain what a panic attack is, or to guide you through a grounding exercise, and it will do a creditable job — the raw material of psychoeducation is exactly the kind of well-documented knowledge these models absorbed in training. A PRISMA review by Guo and colleagues (2024, JMIR Mental Health) of 40 studies found that a large share of the research focuses precisely on these detection-and-conversation tasks, and that LLMs perform them with surprising competence.

There's even evidence that AI can improve human care. The HAILEY system — an AI-in-the-loop tool that suggests edits to peer supporters rather than replacing them — raised counselors' expressed empathy by 19.6% overall, and by nearly 39% among those who were struggling most (Sharma et al., 2023, Nature Machine Intelligence). Used as a helper, not a stand-in, an LLM can be a real asset. The roadmap paper by Stade and colleagues (2024, npj Mental Health Research, 3:12) captures this middle path: it proposes a staged scale of autonomy for LLMs in behavioral care (from L0 to L5), and recommends transparency about the AI's status, active countermeasures against bias, and — tellingly — "empathy in a limited degree." Ambition matched to caution: the point of the levels is that a model can be trusted with more responsibility only as it earns it, one guardrail at a time, rather than all at once.

Where general-purpose LLMs fail in therapy

The failures are not edge cases; they follow from what a general model is.

  • No crisis protocol. A trained clinician has a procedure for someone expressing suicidal intent. A general chatbot has no reliable, tested escalation path — it may say something soothing, something wrong, or something dangerous, and it cannot call anyone.
  • Sycophancy and stigma. Moore and colleagues (2025, ACM FAccT) showed that LLMs express stigma toward certain mental-health conditions and give unsafe responses, concluding they are unfit to replace clinicians. General models are also tuned to be agreeable, which is the opposite of what therapy sometimes requires — a good therapist challenges you; a sycophantic model validates whatever you bring.
  • No memory of your treatment. A general chat starts fresh each session. It doesn't hold your history, your plan, or last month's breakthrough, so it can't build on anything.
  • Hallucinated advice. LLMs state incorrect things confidently. The FDA's Digital Health Advisory Committee (November 2025) put it bluntly: the "metacognitive limitations of AI create significant risks, including potentially fatal misinformation."
  • Cultural bias. These models are trained predominantly on WEIRD populations — Western, Educated, Industrialized, Rich, Democratic — and research has shown that stereotypes about race, gender, and income persist even in the most advanced systems (Wang et al., 2024). Advice tuned to one cultural context can miss or misread how distress is expressed in another.

The clearest cautionary tale is Tessa, a chatbot deployed by the U.S. National Eating Disorders Association. It began giving weight-loss advice to people with eating disorders — the precise opposite of the therapeutic goal — and was pulled offline. It is a compact illustration of how a tool that sounds helpful can do harm without rigorous testing and guardrails.

It's also worth noting how thin the evidence still is. A scoping review by Hua and colleagues (2025, npj Digital Medicine) of studies on LLMs in mental health found evaluation methods that aren't standardized and a heavy reliance on proprietary models that researchers can't inspect. In other words, even the studies that look encouraging are hard to compare with one another — which is a reason for caution, not confidence, when a model claims to help.

Purpose-built vs. general LLM: comparison table

The point isn't that AI can't help in mental health — it's that how the system is built changes everything.

DimensionGeneral LLM (e.g. ChatGPT)Purpose-built therapy tool
Crisis handlingNo tested escalation pathCrisis detection + routing to human help
Evidence baseGeneral benchmarks, not clinicalDesigned around CBT/therapeutic protocols
MemoryResets each sessionLongitudinal, tracks history and plan
GuardrailsTuned to be agreeableSafety boundaries, transparency about AI status

A general model optimizes for a fluent answer; a purpose-built tool optimizes for a safe and clinically-shaped one. Some of the most promising research pushes even further — multi-agent systems where several specialized agents debate the best strategy before responding, which we cover in multi-agent AI therapists vs. a single chatbot. What separates a responsible product from a raw chatbot is precisely the set of safety mechanisms described in our piece on guardrails for mental-health AI: crisis detection, escalation to a human, and honesty about being an AI.

LLMs are not a replacement for a psychotherapist. But they can be a valuable complement — available anytime, without stigma, without a waiting list — when they're wrapped in the right protections. That is the approach Nearby takes: built-in safety mechanisms, evidence-based structure, and the ability to point someone toward a human specialist in a crisis. The key is transparency about what the system is, and clarity about the limits of what it can do.

FAQ

Is it safe to use ChatGPT as a therapist?

For low-stakes reflection — organizing your thoughts, learning a CBT technique, venting on a hard day — a general LLM can be helpful and is broadly low-risk. It becomes unsafe when the stakes rise: it has no crisis protocol, it can state wrong things confidently, and it's tuned to agree with you rather than challenge you. Treat it as a journaling companion that talks back, not as clinical care, and never as a lifeline in a crisis.

What's the difference between ChatGPT and a purpose-built therapy chatbot?

A general LLM is optimized to produce a fluent answer to anything. A purpose-built therapy tool is designed around clinical protocols, keeps a longitudinal memory of your history, is transparent about being an AI, and — most importantly — has a tested path to detect a crisis and route you to a human. Same underlying technology, very different safety engineering around it.

Can ChatGPT handle a mental health crisis?

No. A general chatbot has no reliable way to recognize a crisis, no protocol to follow, and no ability to summon help. If you are thinking about harming yourself or someone else, or you're in acute distress, contact your local emergency services or a crisis line immediately, and reach out to a trusted person. Real human help is available and is what a crisis calls for.

Nearby is a support tool that uses evidence-based psychology. It does not replace a psychologist, psychotherapist, psychiatrist, or emergency service.