paper-with-me

홈 › Papers

VIBE: Video-Input Brain Encoder for fMRI Response Modeling

2025-07-23 · Daniel Carlström Schad, Shrey Dixit, Janis Keck, Viktor Studenyak, Aleksandr Shpilevoi, Andrej Bicanski arxiv

We present VIBE, a two-stage Transformer that fuses multi-modal video, audio, and text features to predict fMRI activity. Representations from open-source models (Qwen2.5, BEATs, Whisper, SlowFast, V-JEPA) are merged by a modality-fusion transformer and temporally decoded by a prediction transformer with rotary embeddings. Trained on 65 hours of movie data from the CNeuroMod dataset and ensembled across 20 seeds, VIBE attains mean parcel-wise Pearson correlations of 0.3225 on in-distribution Friends S07 and 0.2125 on six out-of-distribution films. An earlier iteration of the same architecture obtained 0.3198 and 0.2096, respectively, winning Phase-1 and placing second overall in the Algonauts 2025 Challenge.

📄 PDF Abstract BibTeX arXiv:2507.17958

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

LLM4Brain: Training a Large Language Model for Brain Video Understanding

2024-09-26 · Ruizhe Zheng, Lichao Sun

Decoding visual-semantic information from brain signals, such as functional MRI (fMRI), across different subjects poses significant challenges, including low signal-to-noise ratio, limited data availability, and cross-su…

Domain AdaptationLanguage ModelingLanguage ModellingLarge Language Model+1

CineBrain: A Large-Scale Multi-Modal Brain Dataset During Naturalistic Audiovisual Narrative Processing

2025-03-10 · Jianxiong Gao, Yichang Liu, Baofeng Yang, Jianfeng Feng 외

In this paper, we introduce CineBrain, the first large-scale dataset featuring simultaneous EEG and fMRI recordings during dynamic audiovisual stimulation. Recognizing the complementary strengths of EEG's high temporal r…

DecoderEEGVideo Reconstruction

Stacked Regression using Off-the-shelf, Stimulus-tuned and Fine-tuned Neural Networks for Predicting fMRI Brain Responses to Movies (Algonauts 2025 Report)

2025-10-02 · Robert Scholz, Kunal Bagga, Christine Ahrends, Carlo Alberto Barbano arxiv

We present our submission to the Algonauts 2025 Challenge, where the goal is to predict fMRI brain responses to movie stimuli. Our approach integrates multimodal representations from large language models, video encoders…

Voxel-Level Brain States Prediction Using Swin Transformer

2025-06-13 · Yifei Sun, Daniel Chahine, Qinghao Wen, Tianming Liu 외

Understanding brain dynamics is important for neuroscience and mental health. Functional magnetic resonance imaging (fMRI) enables the measurement of neural activities through blood-oxygen-level-dependent (BOLD) signals,…

Prediction

Graph Autoencoders for Embedding Learning in Brain Networks and Major Depressive Disorder Identification

2021-07-27 · Fuad Noman, Chee-Ming Ting, Hakmook Kang, Raphael C. -W. Phan 외

Brain functional connectivity (FC) reveals biomarkers for identification of various neuropsychiatric disorders. Recent application of deep neural networks (DNNs) to connectome-based classification mostly relies on tradit…

Functional ConnectivityGraph Embedding