BCIwiki (bciwiki.com) — A large-scale magnetoencephalography (MEG) dataset named LibriBrain100 was uploaded to the arXiv preprint server on August 25, 2026, containing over 100 hours of high-quality data, with approximately 80 hours from a single subject, setting a new record for within-subject data depth. The dataset aims to provide a standardized benchmark for neural speech decoding research, accompanied by open-source tools and an online competition.
Using an existing decoding model, the team achieved state-of-the-art performance on a word-classification benchmark, validating data quality and the value of large-scale within-subject data. Additionally, the team collected approximately 40 minutes of extra data from each of 32 subjects, demonstrating the role of multi-subject data in compensating for limited per-subject data. The dataset provides standard train, validation, and test splits, accessible via an open-source Python library for downloading and preprocessing.
The study was conducted by Francesco Mantegna, Dulhan Jayalath, Gereon Elvers, and colleagues, who hope LibriBrain100 will accelerate progress toward non-invasive brain-computer interfaces, potentially restoring communication for people with severe paralysis.