Silent reading as a stand-in — filling a gap in brain data that cannot be collected
Decoding inner speech non-invasively runs into a basic data problem: a corpus pairing brain activity with spontaneous inner monologue cannot be gathered. This work treats silent reading as a scalable stand-in and asks how much can be recovered.
Paper overview (our summary)
- Field (arXiv category)cs.LG(+1)
- AuthorsIngo Marquardt, Anthilia Alchanat, Priyanka Jain
- Submitted2026-08-20
- arXiv ID2608.20186v1
Key points
- Decoding inner speech without surgery is held back by the impossibility of assembling a corpus that pairs brain activity with a spontaneous inner monologue.
- The authors treat silent reading as a scalable proxy task and ask how much lexical and semantic information a contrastive decoder can extract.
- About 240,000 word presentations were recorded from a single participant across 393 runs, roughly 49 hours, on 19-channel dry-electrode EEG.
- Scored as word-grouped top-10 retrieval, decoding sat reliably above chance and rose log-linearly with the volume of training data, without levelling off.
- Dropping the occipital and posterior-temporal electrodes cost about a third of the word-level gain, and left context tracking where it was.
1Working around data that cannot be gathered
This paper begins not from a method but from a constraint on data. When a person is putting words together in their head unprompted, that content and the accompanying brain activity cannot be recorded as a pair. The proxy paradigms used instead have their own difficulties: slow to acquire, poorly time-locked, and offering no way to check whether the subject did as instructed.
2Silent reading as the stand-in
With silent reading, what was read and when is fixed on the experimenter's side. It is not inner speech, but that is exactly what buys the scale and the precision. The choice of proxy task settles what the analysis can reach.
3The scale
The design is to measure one participant deeply rather than many shallowly: about 49 hours from a single person. Typography was randomised on every trial so that word identity would be partly decorrelated from low-level visual form.
4Not yet saturated
Decoding ran reliably above chance and reached mid-frequency and rare words. What draws attention is that it rose log-linearly with the volume of training data, without levelling off. From this the authors conclude that decoding here is not saturated but limited by data.
Removing occipital and posterior-temporal electrodes cut the word-level gain by roughly a third while leaving context tracking unchanged. Control analyses hold word-level decoding apart both from the tracking of narrative context and from a positional prior of non-neural origin, carried in by the positional embedding of the transformer.
Why it matters
Where the thing one wants to measure cannot in principle be recorded, the design of the study collapses into the choice of proxy task. Silent reading is not inner speech, but because what was read and when is fixed, it buys scale and temporal precision. That performance keeps scaling log-linearly with data says the ceiling in this area is the volume of data rather than the method, which puts the proxy design at the centre of how far the work can go.
FAQ
Why use silent reading as a proxy?
Why randomise the typography?
What does no saturation mean here?
Sources (primary)
Source: arXiv (descriptive metadata is CC0 public domain). Summaries are our own; see arXiv for the original text and PDF.
- arXiv abstract page (original, official)
- PDF (arXiv)
- arXiv ID: 2608.20186