cs.LG q-bio.NC

Silent reading as a stand-in — filling a gap in brain data that cannot be collected

cs.LG Ingo Marquardt, Anthilia Alchanat, Priyanka Jain Aug 2026

Decoding inner speech non-invasively runs into a basic data problem: a corpus pairing brain activity with spontaneous inner monologue cannot be gathered. This work treats silent reading as a scalable stand-in and asks how much can be recovered.

Paper overview (our summary)

  • Field (arXiv category)cs.LG(+1)
  • AuthorsIngo Marquardt, Anthilia Alchanat, Priyanka Jain
  • Submitted2026-08-20
  • arXiv ID2608.20186v1

Key points

  • Decoding inner speech without surgery is held back by the impossibility of assembling a corpus that pairs brain activity with a spontaneous inner monologue.
  • The authors treat silent reading as a scalable proxy task and ask how much lexical and semantic information a contrastive decoder can extract.
  • About 240,000 word presentations were recorded from a single participant across 393 runs, roughly 49 hours, on 19-channel dry-electrode EEG.
  • Scored as word-grouped top-10 retrieval, decoding sat reliably above chance and rose log-linearly with the volume of training data, without levelling off.
  • Dropping the occipital and posterior-temporal electrodes cost about a third of the word-level gain, and left context tracking where it was.

1Working around data that cannot be gathered

This paper begins not from a method but from a constraint on data. When a person is putting words together in their head unprompted, that content and the accompanying brain activity cannot be recorded as a pair. The proxy paradigms used instead have their own difficulties: slow to acquire, poorly time-locked, and offering no way to check whether the subject did as instructed.

2Silent reading as the stand-in

The aspectProxy paradigms so farThis work, silent reading
What the subject doesCued repetitive inner speech, or generative inner speech reported afterwardsReads continuous narrative text silently
Speed of acquisitionSlowScales
Time-lockingPoorEach presentation carries its own moment
Verifying complianceNot possibleThe word presented is known

With silent reading, what was read and when is fixed on the experimenter's side. It is not inner speech, but that is exactly what buys the scale and the precision. The choice of proxy task settles what the analysis can reach.

3The scale

Word presentationsAbout 240,000From a single participant
Runs and hours393 runs, about 49 hoursA densely sampled design
The recording19-channel dry-electrode EEGWords shown in rapid serial visual presentation

The design is to measure one participant deeply rather than many shallowly: about 49 hours from a single person. Typography was randomised on every trial so that word identity would be partly decorrelated from low-level visual form.

4Not yet saturated

Decoding ran reliably above chance and reached mid-frequency and rare words. What draws attention is that it rose log-linearly with the volume of training data, without levelling off. From this the authors conclude that decoding here is not saturated but limited by data.

Removing occipital and posterior-temporal electrodes cut the word-level gain by roughly a third while leaving context tracking unchanged. Control analyses hold word-level decoding apart both from the tracking of narrative context and from a positional prior of non-neural origin, carried in by the positional embedding of the transformer.

Why it matters

Where the thing one wants to measure cannot in principle be recorded, the design of the study collapses into the choice of proxy task. Silent reading is not inner speech, but because what was read and when is fixed, it buys scale and temporal precision. That performance keeps scaling log-linearly with data says the ceiling in this area is the volume of data rather than the method, which puts the proxy design at the centre of how far the work can go.

FAQ

Why use silent reading as a proxy?
Spontaneous inner speech cannot be recorded, and existing proxies are slow and poorly time-locked, whereas silent reading fixes what was read and when.
Why randomise the typography?
To decorrelate word identity in part from low-level visual form.
What does no saturation mean here?
That performance continued to scale log-linearly with training data, indicating the limit is the volume of data rather than the method.

Sources (primary)

Source: arXiv (descriptive metadata is CC0 public domain). Summaries are our own; see arXiv for the original text and PDF.

#AI research#Neuroscience#EEG#arXiv#Machine learning
Disclaimer: This site independently summarizes and classifies information based on official data sources. Always verify the latest and accurate information with the official sources. Content on finance, health, legal, and security is information, not advice. This site is not an official website of the U.S. government.