paper-with-me

홈 › Papers

Controllable Mind Visual Diffusion Model

2023-05-17 · Bohan Zeng, Shanglin Li, Xuhui Liu, Sicheng Gao, XiaoLong Jiang, Xu Tang, Yao Hu, Jianzhuang Liu, Baochang Zhang

Brain signal visualization has emerged as an active research area, serving as a critical interface between the human visual system and computer vision models. Although diffusion models have shown promise in analyzing functional magnetic resonance imaging (fMRI) data, including reconstructing high-quality images consistent with original visual stimuli, their accuracy in extracting semantic and silhouette information from brain signals remains limited. In this regard, we propose a novel approach, referred to as Controllable Mind Visual Diffusion Model (CMVDM). CMVDM extracts semantic and silhouette information from fMRI data using attribute alignment and assistant networks. Additionally, a residual block is incorporated to capture information beyond semantic and silhouette features. We then leverage a control model to fully exploit the extracted information for image synthesis, resulting in generated images that closely resemble the visual stimuli in terms of semantics and silhouette. Through extensive experimentation, we demonstrate that CMVDM outperforms existing state-of-the-art methods both qualitatively and quantitatively.

📄 PDF Abstract BibTeX arXiv:2305.10135

Code (1)

zengbohan0217/cmvdm 공식 구현 pytorch

Tasks

AttributeImage Generationmodel

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Batch Normalization 설명 없음
Residual Connection 설명 없음
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Residual Block Residual Blocks are skip-connection blocks that learn residual functions with reference to the layer inputs, instead of learning unreferenced functions. They were introduced…

Similar Papers 제목 키워드 기반

Tokenization Allows Multimodal Large Language Models to Understand, Generate and Edit Architectural Floor Plans

2026-03-12 · Sizhong Qin, Ramon Elias Weber, Xinzheng Lu arxiv

Architectural floor plan design demands joint reasoning over geometry, semantics, and spatial hierarchy, which remains a major challenge for current AI systems. Although recent diffusion and language models improve visua…

Spatial Reasoning

MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural Diffusion

2023-08-08 · Yizhuo Lu, Changde Du, Qiongyi Zhou, Dianpeng Wang 외

Reconstructing visual stimuli from brain recordings has been a meaningful and challenging task. Especially, the achievement of precise and controllable image reconstruction bears great significance in propelling the prog…

Image Reconstruction

MindDiffuser: Controlled Image Reconstruction from Human Brain Activity with Semantic and Structural Diffusion

2023-03-24 · Yizhuo Lu, Changde Du, Dianpeng Wang, Huiguang He

Reconstructing visual stimuli from measured functional magnetic resonance imaging (fMRI) has been a meaningful and challenging task. Previous studies have successfully achieved reconstructions with structures similar to …

Image Reconstruction

Resonant Minds: Closed-Loop Social Avatars with Theory of Mind

2026-06-04 · Jianxu Shangguan, Jing Xu, Hang Ye, Xiaoxuan Ma 외 arxiv

Creating lifelike digital humans with genuine social intelligence requires unifying cognitive reasoning and multimodal generation within a coherent framework. Current approaches treat these as separate tasks: Large Langu…

multimodal generationVideo Generation

MinD: Unified Visual Imagination and Control via Hierarchical World Models

2025-06-23 · Xiaowei Chi, Kuangzhi Ge, Jiaming Liu, Siyuan Zhou 외

Video generation models (VGMs) offer a promising pathway for unified world modeling in robotics by integrating simulation, prediction, and manipulation. However, their practical application remains limited due to (1) slo…

Video GenerationVideo Prediction