Three components for every self-harm patient instead of three risk groups
- Systematic review and meta-analysis of diagnostic accuracy led from the University of Melbourne, with co-authors at KU Leuven, Oxford and Manchester, published in PLOS Medicine on 11 September 2025 and registered as PROSPERO CRD42024523074. Eight databases were searched to 30 April 2025. Studies qualified if the outcome was suicide, hospital-treated self-harm or a composite of the two, measured in a cohort, case-cohort or case-control design; self-reported outcomes and suicidal ideation alone were excluded. Fifty-three studies met the criteria, covering 35 million records and 249,000 occurrences of suicide and self-harm. Quality was assessed with QUADAS-2 and against the TRIPOD checklist.
- Pooled discrimination and accuracy: area under the receiver operating characteristic curve 0.69 to 0.93, sensitivity 45% to 82%, specificity 91% to 95%, positive likelihood ratios 6.5 to 9.9 and negative likelihood ratios 0.2 to 0.6. Accuracy was extracted at the 95th percentile of predicted risk, the threshold most common in this literature, and where a study reported several follow-up points the longest was taken.
- One-year baseline rates estimated from the cohort studies were 0.7% for suicide, 2.1% for hospital-treated self-harm and 1.6% for the combined outcome. Against those rates the positive predictive values were 6% for suicide, 16% and 17% for self-harm and 9% for the combined outcome; negative predictive values ran from 90% to 99%. Applied outside the included samples at a positive likelihood ratio of 10, the positive predictive value was 0.10% in a general population with an annual rate of 0.01%, 17% for suicide after discharge from treatment for self-harm where the annual rate is 2%, and 66% for repeat self-harm in that same group where the annual rate is 16%.
- The authors conclude that this accuracy is too low for case finding and too low for allocating treatment to a high-risk group, and that management after hospital-treated self-harm should instead contain three components offered to every patient: a needs-based assessment and a response to it, identification of modifiable risk factors with treatment intended to reduce those exposures, and implementation of aftercare interventions already demonstrated to be effective. They describe the diagnostic accuracy of machine learning as similar to that of traditional risk assessment scales. Overall risk of bias was rated low in 3 studies, high in 26 and unclear in 24.
Sensitivity of 45% to 82%. Specificity of 91% to 95%. An area under the curve running from 0.69 to 0.93. Those are the pooled figures for 53 machine learning algorithms built to predict suicide and hospital-treated self-harm, and they describe how well the algorithms separate two groups of people. They describe nothing about what a service does once the separation exists, and the review is written around that second question.
Where the positive predictive value lands
Separation becomes a probability only after the base rate is supplied. The review took one-year baseline rates from the cohort studies it had pooled: 0.7% for suicide, 2.1% for hospital-treated self-harm, 1.6% for the two together. Run through Bayes' rule with the pooled likelihood ratios, those rates give positive predictive values of 6% for suicide, 16% and 17% for self-harm, and 9% for the composite. Negative predictive values ran from 90% to 99%.
The authors then applied a hypothetical algorithm with a positive likelihood ratio of 10 to populations outside the review. In the general population, annual suicide rate 0.01%, the positive predictive value came to 0.10%. For suicide within a year of discharge from treatment for self-harm, annual rate 2%, it came to 17%. For repeat self-harm in that same discharged group, annual rate 16%, it came to 66%. The algorithm did not change between those three lines. The population did.
The threshold moves and the errors change places
Accuracy was read at the 95th percentile of predicted risk, and where a study offered several follow-up points the longest was taken, which yields the most favourable positive predictive value available. Raising a threshold buys specificity and spends sensitivity. One study inside the review shows the trade in whole numbers: patients in the top 5% of risk after a mental health specialty visit accounted for 43% of suicide attempts and 48% of suicides over a 90-day window. The remaining 57% and 52% were below the line at the moment of scoring. In the cohort studies the modest sensitivity means more than half of those who later repeat self-harm or die by suicide sit in the low-risk group at the moment of scoring.
Admission, observation, follow-up
Stratification exists to route people. A high-risk classification is what services use to select patients for admission, close observation, or more urgent and more frequent community follow-up, and a low-risk classification is what takes them out of that queue. The review measures its algorithms against that job and finds them insufficient for it, and insufficient for case finding too.
The second half of the argument leaves the accuracy figures behind. Interventions with trial evidence behind them were not tested on high-risk subgroups. Cognitive behavioural therapy for self-harm has been examined in unselected populations and dialectical behaviour therapy in selected ones, and neither design begins by ranking patients along a risk continuum. The classification is not what makes those interventions available.
Assessment, modifiable factors, aftercare
What the authors put in place of the tiers is offered to everyone treated in hospital for self-harm. An assessment of needs, with a response to what it finds. Identification of the risk factors that can be modified, with treatment aimed at reducing those exposures. Aftercare that has already been shown to work, actually implemented.
Several national guidelines arrived there before the algorithms did, and the review says so: the low accuracy of the older risk scales was among the reasons guidance in several countries stopped recommending stratification for allocating aftercare and recommended a psychosocial assessment in its place. The machine learning literature has landed beside the scales it was built to replace, at a similar diagnostic accuracy, and the guidance it was expected to revise stays where it was.
Trials of cognitive behavioural therapy for self-harm recruited unselected patients, so the high-risk label is not what makes that treatment available.
This is a meta-analysis of 53 heterogeneous studies, not a single external validation, and its own quality assessment rates 3 of them at low risk of bias, 26 at high and 24 at unclear; mean adherence to the TRIPOD checklist was 20 of 28 applicable items. Forty-eight otherwise eligible studies were excluded because they reported too little to extract true and false positives and negatives. Follow-up windows differed across studies, from 30 days to 2 years for most, and outcomes were taken at the longest point each study reported, so the accuracy figures are the most favourable each study allows rather than a common horizon. Baseline prevalence was estimated by meta-analysis of the cohort studies and then applied to case-control likelihood ratios as well. Calibration in the prognostic sense, calibration slopes and plots, was not assessed; the review measures diagnostic accuracy. Pooled values could not be produced for two groups of studies. The review does not report a comparison of routing by algorithm against routing by needs assessment, so the three-component recommendation is the authors' proposal rather than a measured outcome. The review examines prediction and says nothing about means restriction, crisis lines or structured post-attempt interventions, which lie outside its question.