BEND-BCI spans motor, visual, speech and spatial decoding tasks across 16 real or synthetic neural recordings, comparing 23 methods on four axes: held-out prediction, robustness to input perturbation, computational cost and cross-recording latent consistency. Those extra axes frequently changed the rankings, and held-out accuracy did not reliably identify the most robust, most efficient or most cross-recording-consistent models. Simpler baselines held their own against far more heavily parameterised deep neural networks, and in some cases outperformed them.
Diagnostic analyses based on explainable machine learning further linked decoder performance to whether a model used the expected neural features, and showed performance could improve by selecting high-quality training trials. The authors frame BEND-BCI as a way to recast decoder selection from an accuracy leaderboard into a constrained decision over task, resource, representation and diagnostic goals. Some of the recordings used in the benchmark are synthetic, and the paper has not yet been peer reviewed.