The control app was never empty: what 32 anxiety trials measured when they measured digital placebo
- Systematic review and meta-analysis from the University of Osaka: 54 randomised trials with a digital placebo comparator (7092 participants) were reviewed, and the 32 trials using GAD-7, DASS-A or HADS-A (5311 participants) were pooled. The pooled change in the control arm was Hedges g = 0.28 (95% CI 0.18 to 0.38), with I² = 76%.
- The estimate is a within-arm pre-post change computed from the baseline and post-intervention means and standard deviations of the comparator group. It is not a contrast against an untreated group, so natural course, regression to the mean and repeated self-rating sit inside the same number as expectancy.
- The estimate held up under pressure: the Egger test gave p = 0.47, trim-and-fill imputed 2 studies and lowered it to g = 0.24 (95% CI 0.13 to 0.35), low-risk-of-bias trials alone gave g = 0.26 (95% CI 0.15 to 0.36), and excluding trials with imputed means or standard deviations gave g = 0.30 (95% CI 0.20 to 0.40).
- What the control app contained was a statistically significant moderator (subgroup p = 0.02), alongside diagnosis (p = 0.03) and baseline severity (p = 0.02); together with scale choice these explained R² = 58.5% of heterogeneity. One caution: the meta-regression row for the "Removed" comparator prints a coefficient of -0.048 that cannot be reconciled with its own interval (-0.954 to -0.013), standard error (0.240) or p = 0.04, all of which imply roughly -0.48.
A group at the University of Osaka asked the question that trials of digital mental health tools mostly step around: how far does the control app move an anxiety score on its own? Across 32 randomised trials and 5311 participants, the pooled answer was Hedges g = 0.28 (95% CI 0.18 to 0.38). The figure is more modest than the enthusiasm around digital placebo would predict, and it measures something other than what its name suggests.
What the number actually measures
The effect size was computed from the baseline and post-intervention means and standard deviations of the comparator arm alone. That is a within-arm change over the treatment period, not a between-group contrast. Everything capable of moving a self-report anxiety score over several weeks is folded into it: the natural course of an anxious episode, regression to the mean in a cohort selected for elevated scores at screening, the reactivity of completing GAD-7 for the fourth time, and whatever expectancy the control app generated. Calling the whole of it a placebo effect assigns to expectancy a quantity that several processes produced together.
The authors were disciplined where the data allowed. They searched PubMed, Web of Science and Scopus in July 2024; where a trial reported more than one anxiety endpoint they took the one with the smallest placebo response, which biases the pooled figure downward rather than up. Publication bias was not detected by the Egger test (p = 0.47), and the trim-and-fill procedure imputed 2 missing studies, bringing the estimate to g = 0.24 (95% CI 0.13 to 0.35). Restricting to trials at low risk of bias gave g = 0.26 (95% CI 0.15 to 0.36); excluding trials whose means or standard deviations had to be imputed gave g = 0.30 (95% CI 0.20 to 0.40). The estimate is stable across every stress test the authors applied. Its interpretation is where the care is needed.
The comparator is not a constant
The review classifies control apps into four kinds, a taxonomy it adopts from earlier work rather than devising, and that classification is the part worth keeping. "Replaced" swaps the active ingredient for an inert one. "Removed" deletes the active ingredient outright. "Unrelated" substitutes a different, unrelated active ingredient. "Less" delivers a weaker version of the same thing. A trial comparing a therapy app against a stripped-down copy of itself and a trial comparing it against a mood-diary shell are not running the same experiment, and their shared label of "placebo-controlled" hides that difference rather than describing it.
The comparator category was a statistically significant moderator (subgroup p = 0.02), as were diagnosis (p = 0.03) and baseline severity (p = 0.02). In the meta-regression, primary psychiatric patients carried a coefficient of 0.308 (95% CI 0.087 to 0.530, p = 0.01) and low baseline scores a coefficient of -0.222 (95% CI -0.406 to -0.038, p = 0.02); with the choice of scale these accounted for R² = 58.5% of heterogeneity. Note the direction on severity: higher starting scores produced larger apparent placebo responses. That is what regression to the mean predicts. It is also what expectancy predicts. This design cannot separate the two.
One caution about the table itself, since the reader may go to it. The row for the "Removed" comparator prints a coefficient of -0.048 next to a 95% interval of -0.954 to -0.013, a standard error of 0.240 and p = 0.04. Those four values cannot all be true: the interval is centred near -0.48, which is precisely what that standard error and that p value imply, and it is the larger value, not -0.048, that matches the paper's own statement that "Removed" controls produced the smallest response. Two further rows show comparable slips, in the lower limit for age and the standard error for number of groups; every other row is internally consistent. I followed the intervals, which agree with the prose, and would not quote the printed coefficient in that cell.
Reading an app trial in clinic
When a patient arrives with an app that has a randomised trial behind it, the useful question is not whether it beat placebo but what the placebo was. A standardised difference of 0.3 over a hollowed-out copy of the same product means something different from the same difference over an unrelated app. Ask what the control group actually opened, how long they used it, and how often they were asked to rate themselves. Trial registries usually answer the first two; the third is often the largest unmeasured ingredient.
The 0.28 belongs to the trial, not to the patient in front of you. In routine care nobody is randomised, nobody is blinded, and no one is measured weekly by a research assistant. What this meta-analysis does establish is a floor: in trials of interventions that cannot be properly blinded, the control arm moves by roughly a quarter to a third of a standard deviation before the intervention's own effect becomes visible. Any digital product has to clear that before its claim means anything.
The severity finding is the one to say out loud. Patients recruited at high baseline scores improved more in the control arm, and the same asymmetry operates in clinic, where people usually seek help in a bad week. When someone starts an app on Monday and feels steadier a fortnight later, the improvement is real and worth acknowledging. Its cause is not established by the improvement itself. Both statements can be said in the same sentence, and saying them together is what keeps a patient from abandoning a working treatment on the strength of a good fortnight, or crediting an app that happened to be present.
The digital placebo effect reported here is a within-arm change in the control group, which means expectancy is one of several processes packed into a single number.
The authors list their own: only adults aged 18 years and older were included, only three self-report anxiety scales were pooled, three databases were searched, there are no long-term data, and the definition of digital placebo was one of several defensible ones. Heterogeneity was high at I² = 76%. One caveat here is mine and not theirs: the design estimates how much control arms change, not how much of that change expectancy caused, and the paper does not raise that distinction, calling the quantity the digital placebo effect throughout. The first author is an employee of NS Pharma, which the paper states does not offer digital health solutions, and states that the research was conducted independently of the company.