paper-with-me

Papers

BrainCLIP: Bridging Brain and Visual-Linguistic Representation Via CLIP for Generic Natural Visual Stimulus Decoding

2023-02-25 · Yulong Liu, Yongqiang Ma, Wei Zhou, Guibo Zhu, Nanning Zheng

Due to the lack of paired samples and the low signal-to-noise ratio of functional MRI (fMRI) signals, reconstructing perceived natural images or decoding their semantic contents from fMRI data are challenging tasks. In this work, we propose, for the first time, a task-agnostic fMRI-based brain decoding model, BrainCLIP, which leverages CLIP's cross-modal generalization ability to bridge the modality gap between brain activity, image, and text. Our experiments demonstrate that CLIP can act as a pivot for generic brain decoding tasks, including zero-shot visual categories decoding, fMRI-image/text matching, and fMRI-to-image generation. Specifically, BrainCLIP aims to train a mapping network that transforms fMRI patterns into a well-aligned CLIP embedding space by combining visual and textual supervision. Our experiments show that this combination can boost the decoding model's performance on certain tasks like fMRI-text matching and fMRI-to-image generation. On the zero-shot visual category decoding task, BrainCLIP achieves significantly better performance than BraVL, a recently proposed multi-modal method specifically designed for this task. BrainCLIP can also reconstruct visual stimuli with high semantic fidelity and establishes a new state-of-the-art for fMRI-based natural image reconstruction in terms of high-level semantic features.

📄 PDF Abstract BibTeX arXiv:2302.12971

Code (1)

YulongBonjour/BrainCLIP 공식 구현 pytorch

Tasks

Brain DecodingImage GenerationImage ReconstructionImage-text matchingText Matching

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Linguistics and Human Brain: A Perspective of Computational Neuroscience

2026-02-09 · Fudong Zhang, Bo Chai, Yujie Wu, Wai Ting Siok 외 arxiv

Elucidating the language-brain relationship requires bridging the methodological gap between the abstract theoretical frameworks of linguistics and the empirical neural data of neuroscience. Serving as an interdisciplina…

Decoding Visual Neural Representations by Multimodal Learning of Brain-Visual-Linguistic Features

2022-10-13 · Changde Du, Kaicheng Fu, Jinpeng Li, Huiguang He

Decoding human visual neural representations is a challenging task with great scientific significance in revealing vision-processing mechanisms and developing brain-like intelligent machines. Most existing methods are di…

Achieving More Human Brain-Like Vision via Human EEG Representational Alignment

2024-01-30 · Zitong Lu, Yile Wang, Julie D. Golomb

Despite advancements in artificial intelligence, object recognition models still lag behind in emulating visual information processing in human brains. Recent studies have highlighted the potential of using neural data t…

Adversarial RobustnessEEGObject Recognition

BrainFLORA: Uncovering Brain Concept Representation via Multimodal Neural Embeddings

2025-07-13 · Dongyang Li, Haoyang Qin, Mingyang Wu, Chen Wei 외 arxiv

Understanding how the brain represents visual information is a fundamental challenge in neuroscience and artificial intelligence. While AI-driven decoding of neural data has provided insights into the human visual system…

Visio-Linguistic Brain Encoding

2022-04-18 · COLING 2022 10 · Subba Reddy Oota, Jashn Arora, Vijay Rowtula, Manish Gupta 외

Enabling effective brain-computer interfaces requires understanding how the human brain encodes stimuli across modalities such as visual, language (or text), etc. Brain encoding aims at constructing fMRI brain activity g…