A digital conversation available at any hour can appeal to someone traveling, working late, or concerned about recognition when seeking help. But availability is not clinical responsibility. A randomized trial of an AI therapy chatbot produced promising symptom findings while leaving a crucial question untested: how would it compare with a qualified therapist delivering appropriate treatment?
Key findings
210 adults: participants randomized in the 2025 Therabot trial, with 106 receiving access and 104 assigned to a waiting list.
Four weeks: the intervention period, with further assessment at eight weeks.
6.13 versus 2.63 points: average depression-score reductions at four weeks in the relevant symptom group.
No human-therapy arm: the study did not establish superiority or equivalence to a course of clinician-delivered therapy. [1]
The waiting-list group improved too
The chart shows reported means, not confidence intervals or the percentage recovered. The original paper contains the statistical comparison. Source: Heinz and colleagues, NEJM AI. [1]
The trial also reported greater symptom reductions for anxiety and eating-related concerns. Those outcomes use different measures and should not be put on one common severity scale.
The defensible conclusion concerns this particular intervention under the study conditions. It is not a general finding about all AI conversations.
Why the comparator changes the headline
A waiting-list comparison asks whether access to the intervention does better than waiting under the trial’s conditions. A head-to-head comparison with treatment asks a different question.
To claim equivalence to human therapy, a study would need a defined active treatment and an appropriate design. Borrowing a percentage from another therapy trial cannot supply the missing comparison.
Recruitment, severity, contact, and follow-up can differ between studies. Those differences remain important even when the resulting numbers look easy to compare.
Feeling understood is not a complete safety assessment
A positive relationship rating can matter to a user. It does not establish that a system recognizes a complex presentation, responds appropriately to deterioration, or takes responsibility for treatment.
Engagement, experience, symptom change, and safety should therefore be reported separately. Time spent chatting is not automatically clinical progress.
For a person with a demanding role, a frictionless conversation may be attractive. The evaluation should ask whether it supports the intended goal without delaying appropriate professional assessment.
The evidence applies to the studied system
Therabot was purpose-built and fine-tuned for the intervention. Dartmouth’s reporting emphasized clinical oversight and further safety work. Its July 2026 recognition of the research was not a second independent efficacy trial. [2] [3]
A similar interface does not make another product clinically equivalent. Ask whether the service being offered is the one evaluated, or how the evidence supports a material change.
The version matters. Research conducted on one system should not become an unlimited clinical claim for every later version or general-purpose chatbot.
A stronger claim requires stronger evidence
| Claim | Evidence needed |
|---|---|
| Better than waiting-list access | A suitable randomized comparison with stated population and period |
| As effective as human therapy | A credible active-comparator equivalence or noninferiority study |
| Safe without oversight | Direct evaluation of autonomous use, escalation, and harms |
| Suitable for complex patients | Evidence in a relevant clinical population |
| Improves long-term recovery | Longer follow-up, functioning, subsequent care, and missing-data reporting |
This table proposes an evidence standard. It does not certify any product or report THE BALANCE outcomes.
Privacy needs answers beyond a reassuring label
WHO’s AI guidance identifies risks involving inaccurate outputs, bias, privacy, cybersecurity, and excessive reliance on automated recommendations. It calls for oversight and evaluation. [4]
For someone with a public profile, ask what information is collected, who can access it, and whether it is used for further system development. A statement that a conversation is private should not replace understandable terms.
Privacy should not imply isolation from appropriate help. A service needs a clear explanation of what happens when the user’s needs exceed the tool’s role.
This article does not evaluate any product’s emergency capability. A conversational tool should not replace emergency services or direct clinical help when urgent assessment is needed.
Portability is not the same as continuity
A phone travels more easily than a clinical team. That is a practical advantage, but it does not answer who coordinates care across locations.
A patient could use a digital support tool while seeing several professionals. The service should explain whether the tool provides education, self-management, an adjunct to treatment, or a treatment pathway.
With appropriate consent, relevant information may need to reach the treating clinician. The patient should also know what to do if a digital response conflicts with an existing plan. Availability alone cannot resolve clinical responsibility.
Convenience should not decide who is suitable
The study does not establish suitability for every child, person in crisis, or patient with severe co-occurring conditions. Eligibility and exclusions determine how widely a finding applies.
A founder’s ability to work independently is not evidence that self-service treatment is clinically appropriate. Financial resources and digital confidence do not replace an assessment of need.
For private providers integrating tools, the question is how the technology fits the actual care plan. A system should not acquire an undefined clinical role simply because it is convenient to make available.
The full cost is not just the subscription
A low price does not automatically mean a lower cost per improved patient. Oversight, escalation, continued use, and effects on other care belong in that comparison.
A meaningful evaluation would define whose costs are counted and which outcomes matter. It would also examine whether people stop using the tool or obtain additional treatment.
No cost-effectiveness result is calculated here. The trial’s symptom findings cannot be converted into a financial advantage without the relevant data.
What the next useful study should answer
An active-comparator trial could test the intervention against a credible alternative, with defined contact and clinical responsibilities. Longer follow-up would address functioning, durability, and additional care.
Adverse outcomes should be measured directly. Recruitment and missing observations need to remain visible so that selected volunteers do not become a claim about every prospective user.
It would also help to distinguish the intervention’s content from the experience of immediate attention. Attention can have value, but that does not establish which component produces benefit or how safely it transfers into ordinary care.
The bottom line
The chatbot trial provides a reason to take a carefully developed intervention seriously. It also gives a reason to reject a broader comparison that was not tested. Ask whether the specific service has relevant evidence, clear responsibilities, and a suitable place in care, not simply whether it sounds like a therapist.
For journalists
Key finding: Therabot produced greater reported symptom changes than waiting-list access in a 210-adult randomized trial.
Important caveat: no human-therapy arm was tested, and average score changes are not recovery percentages.
Suggested attribution: THE BALANCE analysis of the published Therabot trial and AI governance guidance.
Methodology and sources
This focused briefing separates the trial from developer commentary and governance guidance. The chart displays magnitudes of reported mean reductions without inventing uncertainty. It makes no claim about current certification or autonomous-treatment safety.


