paper-with-me

Papers

EmoMind: Decoding Affective Captions from Human Brain fMRI

2026-05-16 · Bilal A. Mohammed, Lin Gu, Ruogu Fang arxiv

Decoding visual experience from brain activity has advanced substantially, but current brain-to-text systems largely recover semantic content while discarding affect. Additionally, language models can generate emotional text when prompted with categorical labels, but such labels collapse rich inter-subject variability into coarse discrete bins. We present EmoMind, the first end-to-end pipeline for decoding affective captions directly from fMRI signals. EmoMind first retrieves a semantically grounded neutral scene description from brain-decoded visual features, then rewrites it using a continuous 34-dimensional emotion vector decoded from the same fMRI recording. To control the balance between content preservation and affective expression, we train the rewriter with classifier-free guidance against an identity-preserving null branch, enabling smooth interpolation between semantic fidelity and affective expressivity. We evaluate affective caption generation with a three-axis validation framework spanning subject-specificity, structural geometry, and causal control. We further augment this framework with a synthetic-brain substitution test that probes robustness to the measurement apparatus, and we benchmark each axis against GPT-4 prompted with brain-decoded top-5 emotion labels as a strong discrete baseline. Across two independent emotion fMRI datasets, EmoMind significantly outperforms label-prompted GPT-4 on all three axes, with the largest gains on metrics that require person-specific affective structure rather than population-level emotion aggregation. These results establish continuous brain-decoded affect as a viable control signal for individualized affective caption generation and open new directions for studying individual affective brain organisation.

📄 PDF Abstract BibTeX arXiv:2605.16739

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Brain Captioning: Decoding human brain activity into images and text

2023-05-19 · Matteo Ferrante, Furkan Ozcelik, Tommaso Boccato, Rufin VanRullen 외

Every day, the human brain processes an immense volume of visual information, relying on intricate neural mechanisms to perceive and interpret these stimuli. Recent breakthroughs in functional magnetic resonance imaging …

Brain DecodingDepth EstimationImage CaptioningImage Reconstruction+2

UniBrain: Unify Image Reconstruction and Captioning All in One Diffusion Model from Human Brain Activity

2023-08-14 · Weijian Mai, Zhijun Zhang

Image reconstruction and captioning from brain activity evoked by visual stimuli allow researchers to further understand the connection between the human brain and the visual perception system. While deep generative mode…

AllBrain DecodingImage CaptioningImage Reconstruction

MindSemantix: Deciphering Brain Visual Experiences with a Brain-Language Model

2024-05-29 · Ziqi Ren, Jie Li, Xuetong Xue, Xin Li 외

Deciphering the human visual experience through brain activities captured by fMRI represents a compelling and cutting-edge challenge in the field of neuroscience research. Compared to merely predicting the viewed image i…

Brain DecodingLanguage ModelingLanguage ModellingSelf-Supervised Learning

Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging

2025-05-28 · Runze Xia, Shuo Feng, Renzhi Wang, Congchi Yin 외

Brain-to-Image reconstruction aims to recover visual stimuli perceived by humans from brain activity. However, the reconstructed visual stimuli often missing details and semantic inconsistencies, which may be attributed …

Image ReconstructionLanguage ModelingLanguage ModellingSemantic Similarity+1

Brain-language fusion enables interactive neural readout and in-silico experimentation

2025-09-28 · Victoria Bosch, Daniel Anthes, Adrien Doerig, Sushrut Thorat 외 arxiv

Large language models (LLMs) have revolutionized human-machine interaction, and have been extended by embedding diverse modalities such as images into a shared language space. Yet, neural decoding has remained constraine…

Zero-shot Generalization