Before Prediction Error Becomes a Biomarker: A Moscow Lab Reads the Fine Print
- In 36 healthy adults, EEG captured three prediction-error signatures (early mismatch negativity, mismatch response, P300) across three auditory paradigms that differed only in how the sounds were arranged in time.
- When a frequency difference was hard to detect, temporal structure was decisive: grouping tones into predictable five-tone bundles raised accuracy above isolated "oddball" tones and tone pairs (hit rate chi-square 16.00, p<0.001; sensitivity d' chi-square 8.85, p=0.012) and enlarged the P300.
- The gain came without awareness: listeners rated all three arrangements as equally difficult, a clean dissociation between what the brain did and what the person noticed.
- Only 21 of 36 listeners benefited most from the predictable bundles; 9 did best with pairs and 6 with single tones, and when attention was withdrawn the automatic error signal appeared only for easy deviations.
Predictive processing has become the working grammar of cognitive neuroscience, but a framework is only as useful as the signals that operationalise it. Mismatch negativity – the brain's automatic response to a violated expectation – is the field's workhorse readout of prediction error and one of psychiatry's most replicated biomarkers. A new EEG study from two Moscow laboratories asks a deceptively plain question: does that signal depend on how you arrange the sounds in time? The answer, unhelpfully for anyone hoping for a plug-and-play biomarker, is yes.
The instruments, not the illness
The team – Krystsina Liaukovich and Olga Martynova, working across the Institute of Higher Nervous Activity and Neurophysiology of the Russian Academy of Sciences and the HSE University Centre for Cognition and Decision Making in Moscow – recorded EEG from 36 healthy adults (mean age 24, range 19–44, 10 men, no formal musical training). Each listener heard the same closely spaced tones (a standard around 480/960/1440 Hz and deviants differing by roughly 1% or 8%) under three arrangements: isolated oddball tones, tone pairs, and predictable five-tone bundles. They did this twice – once attending to the sounds, once ignoring them – which let the authors separate three prediction-error signatures: the early mismatch negativity (a sensory-level error, roughly 50–100 ms), the mismatch response (an involuntary grab of attention, 150–300 ms), and the P300 (attention-driven context updating, 250–600 ms).
What temporal context did
When the frequency difference was easy, arrangement barely mattered – a ceiling. When it was hard, the temporal scaffold changed everything. Grouping the tones into predictable bundles lifted hit rate (chi-square 16.00, p<0.001) and sensitivity d' (chi-square 8.85, p=0.012) above both isolated tones and pairs, and enlarged the P300 – in line with the authors' own reading: the more precise the prediction, in both what and when, the larger and earlier the error signal. Two caveats sharpen the point rather than blunt it. First, the gain was invisible to the listener: subjective difficulty ratings were identical across arrangements, a clean split between neural performance and metacognition. Second, when attention was withdrawn the automatic error response survived only for easy deviations, and only 21 of 36 people showed the bundle advantage at all – 9 did best with pairs, 6 with single tones.
Why the field should care
This is basic auditory science, not a clinical trial, and the honest industry story is infrastructural. The whole clinical promise of mismatch negativity – blunted in schizophrenia, tracked as a psychosis-risk marker, probed in disorders of consciousness – rests on the assumption that the signal means one stable thing. This Moscow work shows the readout is contingent: it moves with temporal structure, with attention, and with the individual. That is precisely the unglamorous standardisation labour a maturing biomarker needs, and it is worth noting that a research community with a century-old neurophysiology lineage (the Institute traces back to the Pavlovian tradition) is doing it in English-language, internationally indexed venues. For clinicians the transferable lesson is narrower and more useful than any headline: if prediction-error measures ever reach the clinic, the fine print of the paradigm – timing, attention, and who is being tested – will not be a nuisance to control away. It will be the measurement.
Before prediction error can serve as a biomarker, someone has to prove the signal means the same thing twice, and this study shows how easily it does not.
The sample is small (n=36) and entirely healthy young adults, so clinical inference is indirect; the three paradigms differ in more than temporal structure, and the authors themselves caution against strong causal claims and call for replication with counterbalanced designs.