BCIwiki (bciwiki.com) — A training framework called MOJO keeps neural decoders accurate when labelled data are scarce, outperforming purely supervised models even in few-shot finetuning, according to a preprint posted on arXiv on July 15, 2026, by Ximeng Mao, Nanda H. Krishna, Avery Hee-Woon Ryoo, Matthew G. Perich, and Guillaume Lajoie. The work has not yet been peer reviewed.
MOJO (Masked autOencoder-based JOint training) combines self-supervised learning via masked autoencoding with supervised objectives for spike-tokenizing neural decoders. The researchers evaluated MOJO on three spiking datasets spanning monkey motor cortex during reaching tasks and multi-regional mouse recordings during vision and decision-making tasks, demonstrating superior performance over purely supervised-trained models. The improvement was especially pronounced when training with limited labelled data, particularly in few-shot finetuning where only a small amount of labelled data from a new session is available.
Incorporating self-supervised learning also yielded more interpretable neuronal representations, improving performance on brain region classification and spike-statistics prediction without explicit optimization for these tasks. The study further showed that MOJO generalizes beyond spiking data to human electrocorticography during speech, where it continued to outperform purely supervised models and achieved performance comparable to neuro-foundation models designed specifically for continuous signals.