paper-with-me

홈 › Papers

Revealing Vision-Language Integration in the Brain with Multimodal Networks

2024-06-20 · Vighnesh Subramaniam, Colin Conwell, Christopher Wang, Gabriel Kreiman, Boris Katz, Ignacio Cases, Andrei Barbu

We use (multi)modal deep neural networks (DNNs) to probe for sites of multimodal integration in the human brain by predicting stereoencephalography (SEEG) recordings taken while human subjects watched movies. We operationalize sites of multimodal integration as regions where a multimodal vision-language model predicts recordings better than unimodal language, unimodal vision, or linearly-integrated language-vision models. Our target DNN models span different architectures (e.g., convolutional networks and transformers) and multimodal training techniques (e.g., cross-attention and contrastive learning). As a key enabling step, we first demonstrate that trained vision and language models systematically outperform their randomly initialized counterparts in their ability to predict SEEG signals. We then compare unimodal and multimodal models against one another. Because our target DNN models often have different architectures, number of parameters, and training sets (possibly obscuring those differences attributable to integration), we carry out a controlled comparison of two models (SLIP and SimCLR), which keep all of these attributes the same aside from input modality. Using this approach, we identify a sizable number of neural sites (on average 141 out of 1090 total sites or 12.94%) and brain regions where multimodal integration seems to occur. Additionally, we find that among the variants of multimodal training techniques we assess, CLIP-style training is the best suited for downstream prediction of the neural activity in these sites.

📄 PDF Abstract BibTeX arXiv:2406.14481

Code (1)

vsubramaniam851/brain-multimodal 공식 구현 pytorch

Tasks

Contrastive LearningLanguage Modelling

Similar Papers 제목 키워드 기반

Modelling Multimodal Integration in Human Concept Processing with Vision-and-Language Models

2024-07-25 · Anna Bavaresco, Marianne de Heer Kloots, Sandro Pezzelle, Raquel Fernández

Representations from deep neural networks (DNNs) have proven remarkably predictive of neural activity involved in both visual and linguistic processing. Despite these successes, most studies to date concern unimodal DNNs…

Multimodal foundation models are better simulators of the human brain

2022-08-17 · Haoyu Lu, Qiongyi Zhou, Nanyi Fei, Zhiwu Lu 외

Multimodal learning, especially large-scale multimodal pre-training, has developed rapidly over the past few years and led to the greatest advances in artificial intelligence (AI). Despite its effectiveness, understandin…

Vision-Language Integration in Multimodal Video Transformers (Partially) Aligns with the Brain

2023-11-13 · Dota Tianai Dong, Mariya Toneva

Integrating information from multiple modalities is arguably one of the essential prerequisites for grounding artificial intelligence systems with an understanding of the real world. Recent advances in video transformers…

Brain encoding models based on multimodal transformers can transfer across language and vision

2023-05-20 · NeurIPS 2023 11

Encoding models have been used to assess how the human brain represents concepts in language and vision. While language and vision rely on similar concept representations, current encoding models are typically trained an…

BrainJanus: A Unified Model for Understanding and Generation across Brain, Vision, and Language

2026-06-29 · Haitao Wu, Qirui Zhang, Zhouheng Yao, Shangquan Sun 외 hf

Modeling the bidirectional correspondence between external sensory stimuli and internal neural activity has emerged as a critical frontier in neuroscience. However, existing approaches predominantly treat brain encoding …

Zero-shot Generalization