paper-with-me

홈 › Papers

BrainBench: Benchmarking Large Language Models for Comprehensive EEG Understanding

2026-08-04 · Yangxuan Zhou, Yuning Chen, Chen Wu, Jiquan Wang, Shijian Li, Gang Pan, Sha Zhao arxiv

Electroencephalography (EEG) analysis extends beyond assigning predefined labels to recordings; it requires workflows connecting natural-language instructions, signal processing, quantitative evidence, and scientific interpretation. We term this capability \emph{comprehensive EEG understanding}. Existing evaluations, however, primarily target isolated decoding tasks or system-specific demonstrations, leaving the competence of large language models (LLMs) insufficiently quantified. We introduce \benchmarkname{}, a unified benchmark for comprehensive, instruction-conditioned EEG understanding. It comprises four subsets---Foundational Analysis, Sleep Assessment, Neurocognitive Assessment, and Physiological Integration---covering 17 datasets, \numcases{} tasks, and over \numinstances{} real-data instances. Given an instruction and EEG recordings with optional physiological signals, a system must perform the analysis and produce a scientifically grounded report and, when required, artifacts. Outputs are assessed through numerical, categorical, set, sequence, semantic, and artifact validation. We evaluate \nummodels{} representative LLMs across more than 100K executions under two paradigms: autonomous code execution with CodeAct and structured agentic analysis with BrainAgent. Results vary substantially across models, subsets, difficulty levels, and execution paradigms, showing that EEG competence depends on the model and its operationalization. \benchmarkname{} provides a reproducible testbed for advancing LLM-based EEG understanding. The code and benchmark will be released soon, with evaluation results continuously updated.

📄 PDF Abstract BibTeX arXiv:2608.04156

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

OmniBrainBench: A Comprehensive Multimodal Benchmark for Brain Imaging Analysis Across Multi-stage Clinical Tasks

2025-11-02 · Zhihao Peng, Cheng Wang, Shengyuan Liu, Zhiying Liang 외 arxiv

Brain imaging analysis is crucial for diagnosing and treating brain disorders, and multimodal large language models (MLLMs) are increasingly supporting it. However, current brain imaging visual question-answering (VQA) b…

BrainBench: Exposing the Commonsense Reasoning Gap in Large Language Models

2026-03-16 · Yuzhe Tang arxiv

Large language models (LLMs) achieve impressive scores on standard benchmarks yet routinely fail questions that any human would answer correctly in seconds. We introduce BrainBench, a benchmark of 100 brainteaser questio…

BrainBench: A Brain-Image Test Suite for Distributional Semantic Models

2016-11-01 · EMNLP 2016 11 · Haoyan Xu, Brian Murphy, Alona Fyshe

How Different AI Chatbots Behave? Benchmarking Large Language Models in Behavioral Economics Games

2024-12-16 · Yutong Xie, Yiyao Liu, Zhuang Ma, Lin Shi 외

The deployment of large language models (LLMs) in diverse applications requires a thorough understanding of their decision-making strategies and behavioral patterns. As a supplement to a recent study on the behavioral Tu…

BenchmarkingChatbotDecision MakingNavigate

Large language models surpass human experts in predicting neuroscience results

2024-03-04 · Xiaoliang Luo, Akilles Rechardt, Guangzhi Sun, Kevin K. Nejad 외

Scientific discoveries often hinge on synthesizing decades of research, a task that potentially outstrips human information processing capacities. Large language models (LLMs) offer a solution. LLMs trained on the vast s…