/ EN
2026-08-23 00:00 Papers Foundations & Methods Translated from EN

Arctop Unveils RLbF, Which Trains LLMs on Real-Time EEG Feedback

Summary Arctop has published a companion article to its paper introducing Reinforcement Learning from Brain Feedback (RLbF), a framework that decodes real-time EEG into cognitive states such as workload and stress and uses them as reward signals to train large language models. Unlike RLHF, which depends on sparse, subjective feedback given after the fact, RLbF supplies continuous, involuntary signals that let a model sense in real time how its words land in a listener's brain, which the article says improves communication. It is the first use of brain signals to train a language model and is already running in Arctop's Isaac app, though it adapts on a single dimension, cognitive workload, and technical details are not fully public.
Why it matters RLbF moves the alignment reward from what users report afterward to what their brains do mid-conversation, a real shift in where training signal comes from, though the evidence so far is a company-published article and a single-dimension deployment rather than peer-reviewed benchmarks.

Arctop发布RLbF框架:用脑电信号训练大语言模型
Image: Arctop

BCIwiki (bciwiki.com) — Arctop published a companion article to its paper introducing 'Reinforcement Learning from Brain Feedback (RLbF)', a framework that uses real-time EEG measurements decoded into cognitive states (e.g., workload, stress) as reward signals to train large language models. The article, posted on Arctop's website on August 23, 2026, serves as a plain-language companion to the paper, which is available on a preprint server.

The article notes that traditional Reinforcement Learning from Human Feedback (RLHF) relies on sparse, subjective, and slow feedback provided by users after the fact, whereas RLbF leverages EEG signals to provide continuous, involuntary feedback, enabling models to perceive in real time how their words affect the listener's brain. The framework defines a five-dimensional cognitive state, including enjoyment, cognitive workload, auditory focus, flow state, and stress, each measured by proprietary models built by Arctop.

Arctop says the framework is already deployed in its app 'Isaac', which currently adapts on a single dimension: cognitive workload. Isaac monitors the user's cognitive state via EEG and adjusts the conversation style accordingly. The article outlines a three-phase fine-tuning path, from supervised fine-tuning to reinforcement learning, and discusses potential risks such as reward hacking.

Compiled by BCIwiki from public sources

Sources · 2
arctop.com 2026-08-23
arctop.com 2026-08-23
Read original ↗
Suggest a correction Revisions · none / EN
© 2026 BCIwiki.com Digest Topics Tips Subscribe Revisions About