paper-with-me

홈 › Papers

MIRAGE: Robust multi-modal architectures translate fMRI-to-image models from vision to mental imagery

2026-05-16 · Reese Kneeland, Cesar Kadir Torrico Villanueva, Jordyn Ojeda, Shuhb Khanna, Jonathan Xu, Paul S. Scotti, Thomas Naselaris arxiv

To be useful for downstream applications, vision decoding models that are trained to reconstruct seen images from human brain activity must be able to generalize to internally generated visual representations, i.e., mental images. In an analysis of the recently released NSD-Imagery dataset, we demonstrated that while some modern vision decoders can perform quite well on mental image reconstruction, some fail, and that state-of-the-art (SOTA) performance on seen image reconstruction is no guarantee of SOTA performance on mental image reconstruction. Motivated by these findings, we developed MIRAGE, a method explicitly designed to train on vision datasets and cross-decode mental images from brain activity. MIRAGE employs a linear backbone and multi-modal text and image features as input to a diffusion model. Feature metrics and human raters establish MIRAGE as SOTA for mental image reconstruction on the NSD-Imagery benchmark. With ablation analysis we show that mental image reconstruction works best when decoders use image features with relatively few dimensions and include guidance from text-based and both high- and low-level image-based features. Our work indicates that--given the right architecture--existing large-scale datasets using external stimuli are viable training data for decoding mental images, and warrant optimism about the future success and utility of mental image reconstruction.

📄 PDF Abstract BibTeX arXiv:2605.17198

Code (0)

등록된 구현이 없습니다.

Tasks

Image Reconstruction

Similar Papers 제목 키워드 기반

MIRAGE: Adaptive Multimodal Gating for Whole-Brain fMRI Encoding

2026-05-28 · Abdulkadir Gokce, Badr AlKhamissi, Martin Schrimpf arxiv

Recent progress in task-optimized neural networks has established encoding models as a powerful tool for predicting brain responses to naturalistic stimuli, yet most existing approaches rely on unimodal representations. …

Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation

2025-10-28 · Alexander Martin, William Walden, Reno Kriz, Dengjia Zhang 외 arxiv

We introduce MiRAGE, an evaluation framework for retrieval-augmented generation (RAG) from multimodal sources. As audiovisual media becomes a prevalent source of information online, it is essential for RAG systems to int…

fMRI-LM: Towards a Universal Foundation Model for Language-Aligned fMRI Understanding

2025-11-24 · Yuxiang Wei, Yanteng Zhang, Xi Xiao, Chengxuan Qian 외 arxiv

Recent advances in multimodal large language models (LLMs) have enabled unified reasoning across images, audio, and video, but extending such capability to brain imaging remains largely unexplored. Bridging this gap is e…

Seeing Voices: Generating A-Roll Video from Audio with Mirage

2025-06-09 · Aditi Sundararaman, Amogh Adishesha, Andrew Jaegle, Dan Bigioi 외

From professional filmmaking to user-generated content, creators and consumers have long recognized that the power of video depends on the harmonious integration of what we hear (the video's audio track) with what we see…

Speech Synthesistext-to-speechText to SpeechVideo Generation

MiRAGE: A Multiagent Framework for Generating Multimodal Multihop Question-Answer Dataset for RAG Evaluation

2026-01-21 · Chandan Kumar Sahu, Premith Kumar Chilukuri, Matthew Hetrich arxiv

The rapid evolution of Retrieval-Augmented Generation (RAG) toward multimodal, high-stakes enterprise applications has outpaced the development of domain specific evaluation benchmarks. Existing datasets often rely on ge…

Information RetrievalVisual Grounding