
BCIwiki (bciwiki.com) — Arctop published a companion article to its paper introducing 'Reinforcement Learning from Brain Feedback (RLbF)', a framework that uses real-time EEG measurements decoded into cognitive states (e.g., workload, stress) as reward signals to train large language models. The article, posted on Arctop's website on August 23, 2026, serves as a plain-language companion to the paper, which is available on a preprint server.
The article notes that traditional Reinforcement Learning from Human Feedback (RLHF) relies on sparse, subjective, and slow feedback provided by users after the fact, whereas RLbF leverages EEG signals to provide continuous, involuntary feedback, enabling models to perceive in real time how their words affect the listener's brain. The framework defines a five-dimensional cognitive state, including enjoyment, cognitive workload, auditory focus, flow state, and stress, each measured by proprietary models built by Arctop.
Arctop says the framework is already deployed in its app 'Isaac', which currently adapts on a single dimension: cognitive workload. Isaac monitors the user's cognitive state via EEG and adjusts the conversation style accordingly. The article outlines a three-phase fine-tuning path, from supervised fine-tuning to reinforcement learning, and discusses potential risks such as reward hacking.