Participants underwent MEG recording during auditory and visual imagery tasks, and the signals were source-reconstructed within modality-specific cortical regions of interest. The team compared the CNN and the linear logistic regression model in a subject-specific classification framework. Both approaches decoded above chance, and the CNN beat the linear model in both tasks.
A key finding was cross-modal performance: the CNN still decoded significantly when trained only on cortical regions not relevant to the task. The authors interpret this as evidence that imagined stimuli are represented in distributed, partially overlapping neural networks across modalities rather than in a single sensory area.
The paper says this cross-modal decoding capability highlights the potential of deep learning models to capture complex, multimodal neural patterns, and suggests future brain-computer interfaces could benefit from integrating auditory and visual information, pointing toward more flexible and personalized BCI designs. The findings speak to both cognitive neuroscience and BCI research, though the small sample means they need validation at larger scale.