paper-with-me

홈 › Papers

Aligning Brain Signals with Multimodal Speech and Vision Embeddings

2025-10-29 · Kateryna Shapovalenko, Quentin Auster arxiv

When we hear the word "house", we don't just process sound, we imagine walls, doors, memories. The brain builds meaning through layers, moving from raw acoustics to rich, multimodal associations. Inspired by this, we build on recent work from Meta that aligned EEG signals with averaged wav2vec2 speech embeddings, and ask a deeper question: which layers of pre-trained models best reflect this layered processing in the brain? We compare embeddings from two models: wav2vec2, which encodes sound into language, and CLIP, which maps words to images. Using EEG recorded during natural speech perception, we evaluate how these embeddings align with brain activity using ridge regression and contrastive decoding. We test three strategies: individual layers, progressive concatenation, and progressive summation. The findings suggest that combining multimodal, layer-aware representations may bring us closer to decoding how the brain understands language, not just as sound, but as experience.

📄 PDF Abstract BibTeX arXiv:2511.00065

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mind the Gap: Aligning the Brain with Language Models Requires a Nonlinear and Multimodal Approach

2025-02-18 · Danny Dongyeop Han, Yunju Cho, Jiook Cha, Jay-Yoon Lee

Self-supervised language and audio models effectively predict brain responses to speech. However, traditional prediction models rely on linear mappings from unimodal features, despite the complex integration of auditory …

Prediction

Dynamic Neural Communication: Convergence of Computer Vision and Brain-Computer Interface

2024-11-14 · Ji-Ha Park, Seo-Hyun Lee, Soowon Kim, Seong-Whan Lee

Interpreting human neural signals to decode static speech intentions such as text or images and dynamic speech intentions such as audio or video is showing great potential as an innovative communication tool. Human commu…

Brain Computer Interface

MindAlign: Decoding Inner Speech from fMRI Signals via Multimodal Embedding Alignment under Limited Data

2026-06-15 · Muxuan Liu, Ichiro Kobayashi, Satoshi Nishida arxiv

Decoding inner speech from non-invasive brain signals remains a fundamental challenge due to the absence of overt linguistic output, limited training data, and large inter-subject variability. Existing brain-to-text appr…

Text Generation

A Dataset Generation Scheme Based on Video2EEG-SPGN-Diffusion for SEED-VD

2025-08-30 · Yunfei Guo, Tao Zhang, Wu Huang, Yao Song arxiv

This paper introduces an open-source framework, Video2EEG-SPGN-Diffusion, that leverages the SEED-VD dataset to generate a multimodal dataset of EEG signals conditioned on video stimuli. Additionally, we disclose an engi…

Data Augmentation

MapGuide: A Simple yet Effective Method to Reconstruct Continuous Language from Brain Activities

2024-03-26 · Xinpei Zhao, Jingyuan Sun, Shaonan Wang, Jing Ye 외

Decoding continuous language from brain activity is a formidable yet promising field of research. It is particularly significant for aiding people with speech disabilities to communicate through brain signals. This field…

Text Generation