FDA Convenes First Advisory Committee on Generative AI Therapy Chatbots
- On November 6, 2025, FDA's Digital Health Advisory Committee (DHAC) held its first meeting devoted specifically to generative AI-enabled mental health devices, reviewing a hypothetical prescription LLM chatbot for treating major depressive disorder (MDD) in adults.
- The committee (nine standing members plus temporary experts) evaluated three use cases – diagnosis, treatment, and monitoring – and flagged hallucination, "sycophancy" (a model's tendency to agree with users at the expense of accuracy), bias, and missed clinical deterioration as core risks.
- Members favored a risk-based, total-product-life-cycle framework: staged validation moving from clinician supervision to semi-autonomous use, mandatory definitions of serious adverse events (suicidal ideation, self-harm), a required escalation pathway to a human clinician, and postmarket monitoring of engagement and symptom trajectories.
- The meeting drew 116 public comments by the December 8, 2025 deadline; the committee reached no formal clearance decision, and its recommendations are non-binding.
For the first time, a federal advisory committee sat down specifically to weigh generative AI in psychiatric care – not a wellness app, but a hypothetical prescription chatbot for treating major depressive disorder. The committee's answer was neither a green light nor a ban, but a scaffold: supervise first, prove safety, then loosen the leash.
The Sycophancy Problem
The committee's risk list read like session notes from a difficult supervision hour: hallucinated facts, biased training data, and a subtler failure – sycophancy, the tendency of a model optimized for user satisfaction to validate whatever the patient already believes, even when validation is the wrong move for someone ruminating or catastrophizing. A supervised trainee who drifts into pure validation gets corrected in the next case conference. A chatbot deployed at scale drifts the same way without anyone in the room to notice, and the failure surfaces only after a serious adverse event forces a retrospective look.
A Life-Cycle Model, Not a Yes-or-No
Rather than asking whether to clear such tools, the committee asked how to phase them in: start under clinician supervision, require sponsors to define serious adverse events – suicidal ideation, self-harm – before deployment, build an explicit escalation path to a human when risk indicators appear, and keep monitoring engagement and symptom trajectories after launch instead of treating premarket data as the whole story. None of it is binding: the committee's recommendations carry no regulatory force, and no product has cleared review under this framework – the scenario itself was hypothetical. But the direction is legible. US regulators are drafting guardrails before the market fills the gap, and states are not waiting for federal clarity either – Illinois now requires AI-disclosure in therapy-adjacent chat, California bars chatbots from impersonating licensed therapists.
A supervised trainee who drifts into pure validation gets corrected in the next case conference. A chatbot deployed at scale drifts the same way with nobody in the room to notice.
The meeting produced no formal clearance decision, and the committee's recommendations are not binding on future submissions. No device on the US market has undergone review under this framework – the case discussed was a hypothetical prescription chatbot for MDD, not an existing product. This is a regulatory signal on generative AI oversight generally, not a determination on any specific computational or predictive-processing biomarker.