Megadose AI progress, ranked daily.

LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale

· HF Daily Papers ·
The dataset centers on 100-plus hours of MEG recordings for speech decoding, including about 80 hours from one subject.

LibriBrain100 is built for reproducible benchmarking, with standard train, validation, and test splits and an open Python library for downloading, preprocessing, and loading the data. The authors report state-of-the-art word-classification results using an existing decoding model, arguing that the depth of within-subject data matters. They also include roughly 40 minutes of data from each of 32 additional subjects, showing that supervised finetuning can help when per-user data is limited. The release is paired with an open machine-learning competition and public leaderboard. HF Daily Papers' note

score 5

Categories: Research