ERP-XTTN routes input electroencephalographic peaks to fixed difference-wave prototypes via query-key-only cross-attention with no value projection. Classification is based directly on prototype similarity and a separate measure of component amplitude, so the prototype content contributes to every decision by construction. Prototypes are derived automatically from prominent extrema in the training-fold grand-average difference wave. The team evaluated the architecture on three public datasets (BNCI Horizon 2020, HRI Cursor, and ERP CORE) encompassing eight ERP components (ERN, LRP, ErrP, N170, P300, N2pc, MMN, N400). Evaluations used leave-one-subject-out (LOSO) cross-validation with causal filtering at a three-channel montage, compared against EEGNet, EEG-Deformer, ERP Prototypical Matching Net (EPMN), and xDAWN with Riemannian geometry (xDAWN+RG). At three channels, the mean performance gap between the best baseline and ERP-XTTN was 0.025 area under the receiver operating characteristic curve (AUROC).
Prototype interventions confirmed that decisions depend on the physiological content of the prototypes rather than on the routing attention pattern alone. Across all datasets, false positives morphologically resembled true positives more than true negatives did, indicating that classification errors are neurophysiologically explicable. The researchers noted that, unlike post-hoc explanation methods for black-box models, the basis of each decision is directly observable in the trained model itself. To their knowledge, this is the first epoch-level LOSO benchmark on ERP CORE.