/ EN
2026-09-02 00:00 Papers Foundations & Methods Translated from EN

NVOL: Mid-Layer CLIP Alignment Lifts EEG Image Retrieval to 86.4%

Summary A preprint from Minyi Wang, Zhenqin Wu and Rihui Li reports that aligning EEG signals to an intermediate CLIP layer rather than the final one raised 200-way image retrieval on the THINGS-EEG dataset to 78.1% mean Top-1 accuracy, and to 86.4% with CSLS. Layer-wise contrastive learning selects that layer, which the authors call the Neural Visibility Optimal Layer (NVOL); the same representation then drives generation, with a conditional diffusion prior reconstructing subject-specific NVOL features and mapping them into CLIP space for Stable Diffusion XL, which the authors say beats single-stage final-layer diffusion on semantic and structural metrics. The paper was posted to arXiv on September 2, 2026, and has not been peer reviewed.
Why it matters EEG visual decoding has largely treated CLIP's final layer as the alignment target by default, a choice inherited from vision-language work rather than from anything about neural signals. Letting contrastive learning pick the layer, and then reusing that same representation for reconstruction instead of a separate generative head, is a cheap change with a measurable payoff, which makes it easy for other groups to test. It is a preprint on a single public dataset, so the layer choice may not transfer.

BCIwiki (bciwiki.com) — A preprint study proposes selecting an intermediate CLIP layer as the Neural Visibility Optimal Layer (NVOL) via layer-wise contrastive learning for each subject, which significantly improves EEG-based visual image retrieval accuracy. On the THINGS-EEG dataset, NVOL-based retrieval achieved 78.1% mean Top-1 accuracy in 200-way retrieval, rising to 86.4% with cross-domain similarity local scaling (CSLS). The study by Minyi Wang, Zhenqin Wu, and Rihui Li was posted to arXiv on September 2, 2026 (arXiv:2609.02582) and has not been peer reviewed.

The authors note that most existing visual decoding pipelines directly align EEG features with semantic features from pretrained vision models, but EEG signals carry information at multiple levels, and such practice disregards the varying neural visibility of different visual components, leading to cross-modal mismatches. The framework uses layer-wise contrastive learning to select the intermediate CLIP layer that maximizes retrieval performance as NVOL, and builds a hierarchical framework coupling retrieval and generation on it: the retrieval branch fuses multi-NVOL features, aligns them to image embeddings via contrastive learning, and applies CSLS at test time to mitigate hubness; the generation branch reconstructs subject-specific NVOL features from EEG using a conditional diffusion prior, maps them to CLIP space through a lightweight adapter, and drives a pretrained Stable Diffusion XL model. Experiments showed that two-stage NVOL-to-semantic reconstruction outperforms single-stage final-layer diffusion on semantic and structural metrics.

The study was conducted by Minyi Wang, Zhenqin Wu, and Rihui Li.

Compiled by BCIwiki from public sources

Sources · 1
arxiv.org 2026-09-02
Read original ↗
Suggest a correction Revisions · none / EN
© 2026 BCIwiki.com Digest Topics Tips Subscribe Revisions About