Computational Psychiatry, a Decade In: The Pivot to Trials and the Reliability Problem Underneath It
- In an October 2025 Nature Computational Science comment (vol. 5, pp. 841–843), Quentin Huys (UCL) and Michael Browning (Oxford) argue that computational psychiatry must now move from describing disorders to changing outcomes – delivering causal evidence through clinical trials, drug repurposing, novel interventions and the scaling of psychotherapy.
- The pivot is a quiet concession: a decade after the US NIMH RDoC framework reoriented psychiatry toward mechanism, the field's main yield is a shared vocabulary for how the brain predicts, learns and infers, not deployable clinical tests.
- A 2025 reliability study (n=179, psychosis-spectrum patients and controls) found that computationally derived task parameters had only poor-to-moderate test-retest reliability (ICCs 0.30–0.61) and did not outperform simple behavioural summary scores (ICCs 0.24–0.54).
- Because a measure's reliability caps its power to track symptoms or treatment response, these parameters are not yet stable enough to serve as biomarkers or trial endpoints; the tasks and models need optimising before clinical use.
Every few years psychiatry adopts a framework that promises to replace its symptom checklists with mechanism. Computational psychiatry – the project of modelling mental disorders as specific failures of prediction, learning and inference in the brain – has been the most intellectually seductive of these. A new comment from two of the field's own leaders draws an unusually candid line under the decade, and it is worth reading precisely because it is written by insiders, not critics.
From RDoC to the clinic: a decade-long bet
When the US National Institute of Mental Health launched its Research Domain Criteria (RDoC) around 2010, it bet that biology and computation, not DSM categories, would carve psychiatry at its joints. Computational psychiatry became the sharp edge of that bet. Its work divides into three mathematically related strands: dynamical-systems models, Bayesian inference – the home of predictive processing and Friston's active-inference / free-energy framework – and reinforcement learning. The conceptual payoff has been real: psychosis reframed as aberrant precision-weighting of prior beliefs against sensory prediction errors; depression and anxiety as distortions of reward and punishment learning rates. For the first time, clinicians had a vocabulary that tied a hallucination or an anhedonic slump to a specific, testable computation.
In their October 2025 comment, Huys and Browning argue the field is now "increasingly delivering causal evidence by focusing on interventions research and clinical trials," and that this evidence could improve outcomes through better precision, drug repurposing, novel interventions and the scaling of psychotherapy. Read carefully, that is both an ambition and an admission. Mapping disorders onto computations – back-translation – has been done many times over. Forward translation, from a computation to a treatment decision at the bedside, has remained rare.
The reliability reckoning
The reason forward translation stalls is unglamorous: measurement. For a computational parameter to become a biomarker or a trial endpoint, it must at minimum give the same patient roughly the same value twice. In 2025 that assumption was tested directly. Studying 179 adults across the psychosis spectrum, researchers found that parameters extracted from a standard reward-and-reversal learning task had only poor-to-moderate test-retest reliability (ICCs 0.30–0.61) – and, tellingly, did not outperform the plain behavioural scores they were meant to improve upon (ICCs 0.24–0.54). Related work on reinforcement-learning tasks finds that learning-rate parameters can be reasonably stable while others, such as lapse or exploration terms, sit close to noise. Reliability sets a ceiling on everything downstream: a measure that is unstable cannot reliably correlate with symptoms, predict response, or register a drug effect in a trial. Elegant models built on unstable measurements inherit the instability.
What it means for your practice
Nothing here says computational psychiatry has failed; it says the field is entering the phase where claims get audited against clinical arithmetic. Two practical stances follow. First, when a paper, a start-up or a conference talk offers you a "computational biomarker" for diagnosis or treatment matching, treat it as a research instrument, not a validated test – ask about test-retest reliability before you ask about neural mechanism. Second, use the framework for what it already does well: as a shared language for mechanism that sharpens case formulation. Telling a patient that their voices may reflect the brain over-weighting its own predictions is clinically useful today; ordering a task battery to set their dose is not. Watch the move to trials that Huys and Browning describe – that is where the framework will finally prove, or fail to prove, its worth at the bedside.
A model can be mathematically elegant and mechanistically plausible and still be clinically useless if the same patient scores differently next week.
The reliability figures come from specific decision-making tasks over short retest windows and may not generalise to every computational measure or to neuroimaging-based markers; parameters were somewhat more stable in healthy than in clinical samples. The Huys–Browning piece is an agenda-setting comment, not evidence that the turn to trials will succeed.