A Working Alliance With a Chatbot? What Therabot Measured – and What It Did Not
- In the first randomised trial of a generative-AI therapy chatbot (Therabot, Dartmouth; NEJM AI, 2025), 210 adults were randomised to four weeks of Therabot (n=106) or a waitlist control (n=104); on the WAI-SR, participants' self-rated working alliance reached the band reported for human therapists in outpatient psychotherapy.
- Symptom reductions versus waitlist were large: 51% for depression (d≈0.85), 31% for anxiety (d≈0.79), and 19% for eating-disorder risk (d≈0.63); mean use was about 6 hours, roughly eight therapy sessions' worth.
- Over the same four weeks, the study team intervened 28 times: 15 for safety concerns such as suicidal ideation, and 13 to correct inappropriate chatbot responses.
- A working-alliance rating is a self-report of the patient's experience, not a clinical outcome and not a safety measure; the comparator was a waitlist, not a human-therapist arm, and a NEJM AI commentary noted the WAI was built for human relationships.
In March 2025 a Dartmouth team published the first randomised trial of a generative-AI therapy chatbot, and one line traveled further than the rest: patients rated their working alliance with the software about as highly as patients rate human therapists. A year of argument later, that sentence is still carrying more weight than the data license. It is worth separating what the WAI-SR captured from what it left untouched.
What the trial reported
The design was clean. 210 adults with clinically significant depression, anxiety, or high risk for an eating disorder were randomised to four weeks of unlimited Therabot access (n=106) or a waitlist control (n=104), with assessments at weeks 4 and 8. Reductions relative to waitlist were substantial: depression fell 51% (d≈0.85), anxiety 31% (d≈0.79), and body-image and eating concerns 19% (d≈0.63).
Engagement was the striking part. Participants used the bot for roughly six hours over the month – about eight sessions' worth – much of it initiated by users themselves, often at night. On the WAI-SR, the mean alliance rating landed inside the range typically reported for outpatient psychotherapy. Taken together this is not nothing: people talked to the agent, kept returning, and felt better on self-report.
A bond is not an outcome, and not safety
The reading to resist is the slide from "alliance as high as a human's" to "as good as a therapist." There are three seams. First, the comparator was a waitlist, not a therapist and not an active control, so symptom change cannot be attributed to the alliance – or to anything specific – over simply waiting. Second, the alliance figure is a self-report of the felt bond, goals and tasks; a NEJM AI commentary flagged that the WAI was designed to measure human relationships and its validity on a non-human agent is unestablished, so a high score may index engagement and non-judgment as much as a therapeutic bond. Third, and most concrete: in the same four weeks the team had to step in 28 times – 15 for safety issues including suicidal ideation, 13 to correct inappropriate responses. The alliance felt human; the safety envelope did not hold itself. The lead author was blunt that no generative agent is ready to operate autonomously in mental health.
What to take from it cuts both ways. The finding is real and deserves neither inflation nor a shrug. Patients can form a felt working alliance with a conversational agent, and that is worth studying as a mechanism of engagement in its own right. It is simply not evidence that the agent delivers therapy, is safe unsupervised, or can stand in for the relationship it imitates.
A bond a patient rates as high as a human's tells us about the patient's experience, not about whether the software is doing therapy.
A waitlist with no active or attention control prevents attributing symptom change to the alliance; this is a single four-week trial with selection toward younger, AI-open users, no blinded independent evaluation, and no validated basis for applying the WAI to a non-human agent.