paper-with-me

홈 › Papers

ViewMask-1-to-3: Multi-View Consistent Image Generation via Multimodal Discrete Diffusion Models

2025-12-16 · Ruishu Zhu, Zhihao Huang, Jiacheng Sun, Ping Luo, Hongyuan Zhang, Xuelong Li arxiv

Motivated by discrete diffusion's success in language-vision modeling, we explore its potential for multi-view generation, a task dominated by continuous approaches. We introduce ViewMask-1-to-3, formulating multi-view generation as a discrete sequence modeling problem where each viewpoint is represented as visual tokens from MAGVIT-v2. Through discrete diffusion via masked token prediction, our approach enables progressive multi-view generation via iterative token unmasking, unifying language and vision in a shared token space. Importantly, simple random masking combined with self-attention naturally encourages cross-view consistency without specialized architectures or 3D geometric priors. Our method outperforms the baseline on the GSO and 3D-FUTURE benchmarks, ranking first on average across standard image metrics, and achieving a 10.6% higher IoU than continuous diffusion models on 3D-FUTURE. Furthermore, the proposed framework can be naturally extended to support text-to-image generation and multimodal understanding, highlighting its potential toward a more unified paradigm for multimodal understanding and generation.

📄 PDF Abstract BibTeX arXiv:2512.14099

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving

2025-07-01 · Zeming Chen, Hang Zhao arxiv

Multi-view image generation in autonomous driving demands consistent 3D scene understanding across camera views. Most existing methods treat this problem as a 2D image set generation task, lacking explicit 3D modeling. H…

Scene UnderstandingAutonomous DrivingScene GenerationImage Generation

SyncDreamer: Generating Multiview-consistent Images from a Single-view Image

2023-09-07 · YuAn Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long 외

In this paper, we present a novel diffusion model called that generates multiview-consistent images from a single-view image. Using pretrained large-scale 2D diffusion models, recent work Zero123 demonstrates the ability…

3D GenerationImage to 3DNovel View SynthesisSingle-View 3D Reconstruction+1

MVHuman: Tailoring 2D Diffusion with Multi-view Sampling For Realistic 3D Human Generation

2023-12-15 · Suyi Jiang, Haimin Luo, Haoran Jiang, Ziyu Wang 외

Recent months have witnessed rapid progress in 3D generation based on diffusion models. Most advances require fine-tuning existing 2D Stable Diffsuions into multi-view settings or tedious distilling operations and hence …

3D GenerationDenoising

C^2M-DoT: Cross-modal consistent multi-view medical report generation with domain transfer network

2023-10-09 · Ruizhi Wang, Xiangtao Wang, Jie zhou, Thomas Lukasiewicz 외

In clinical scenarios, multiple medical images with different views are usually generated simultaneously, and these images have high semantic consistency. However, most existing medical report generation methods only con…

Contrastive LearningMedical Report Generation

MVD-Fusion: Single-view 3D via Depth-consistent Multi-view Generation

2024-04-04 · CVPR 2024 1 · Hanzhe Hu, Zhizhuo Zhou, Varun Jampani, Shubham Tulsiani

We present MVD-Fusion: a method for single-view 3D inference via generative modeling of multi-view-consistent RGB-D images. While recent methods pursuing 3D inference advocate learning novel-view generative models, these…

DenoisingDepth EstimationDepth Prediction