De-identified dream reports can be linked across time
A computer often linked a person’s later DreamBank reports to their earlier reports using only a fixed list of common grammar and discourse words. Names, places, occupations, topic terms, and report length were removed as clues.
What the investigation tested
Can grammar reconnect a record to its contributor?
We used 20 individual DreamBank contributors. The model learned from the first 80% of each person’s ordered reports, then tried to link the final 20% back to the correct earlier contributor record.
The privacy mask retained only 161 common grammar and discourse words, such as and, because, with, and although. Names, nouns, places, occupations, and topic terms were excluded. Each report was also normalized so simple verbosity could not drive the result.
Balanced accuracy was 31.9%, compared with 5% chance. The correct contributor appeared among the first three matches 51.0% of the time, compared with 15% chance. Both permutation tests gave p=0.000999.
The result was not universal
Twelve of 20 contributors linked above top-match chance. Eight did not, and six had zero correct first matches. The signal is real at the group level but highly uneven across people.
The stronger idea failed
Waking prose did not reliably identify dream authors
A second frozen test trained on ordinary waking prose from five contributors and tried to identify their dream reports. Grammar reached 25.8% balanced accuracy against 20% chance; scrubbed vocabulary reached 25.5%. Neither survived the correction and effect-size gates.
That null changes the interpretation. The current result looks more like a stable dream-reporting or collection signature than a universal personal fingerprint that crosses from waking thought into dreams.
Scientific value
Why this matters
A privacy warning
Removing names and obvious personal details may not prevent longitudinal narrative records from being linked. Grammar and reporting habits can remain identifying.
A machine-learning warning
Randomly splitting reports by row can let a model learn the contributor instead of the phenomenon. Longitudinal text should be split by person or collection unit.
A boundary on the science
A separate waking-writing test failed. The evidence supports archive linkage, not a universal language fingerprint that passes from waking thought into dreams.
What it does not show
- It does not recover anyone’s legal identity.
- It does not show that dreams reveal biological identity.
- It does not generalize beyond DreamBank yet.
- It does not separate the person from recording or transcription method.
The decisive next test
Use a new cohort with independently collected waking and dream writing, separate recorder and transcriber identities, and a true open-set test that can answer “unknown person.” That would show whether the signal belongs to a person, a recording method, or this archive.
How we checked it
- Both analyses were preregistered and SHA-256 frozen before results.
- Known-answer positive and shuffled-label controls passed.
- Privacy endpoints used 1,000 frozen label permutations.
- A separate sklearn implementation reproduced every headline number to 12 decimal places.
Source: DreamBank, Adam Schneider and G. William Domhoff, UC Santa Cruz. DreamBank data are licensed CC BY-NC-SA 4.0.