paper-with-me

홈 › Papers

Consistent Multimodal Generation via A Unified GAN Framework

2023-07-04 · Zhen Zhu, Yijun Li, Weijie Lyu, Krishna Kumar Singh, Zhixin Shu, Soeren Pirk, Derek Hoiem

We investigate how to generate multimodal image outputs, such as RGB, depth, and surface normals, with a single generative model. The challenge is to produce outputs that are realistic, and also consistent with each other. Our solution builds on the StyleGAN3 architecture, with a shared backbone and modality-specific branches in the last layers of the synthesis network, and we propose per-modality fidelity discriminators and a cross-modality consistency discriminator. In experiments on the Stanford2D3D dataset, we demonstrate realistic and consistent generation of RGB, depth, and normal images. We also show a training recipe to easily extend our pretrained model on a new domain, even with a few pairwise data. We further evaluate the use of synthetically generated RGB and depth pairs for training or fine-tuning depth estimators. Code will be available at https://github.com/jessemelpolio/MultimodalGAN.

📄 PDF Abstract BibTeX arXiv:2307.01425

Code (0)

등록된 구현이 없습니다.

Tasks

multimodal generation

Similar Papers 제목 키워드 기반

Odoriko: A Shape-Aware Multimodal Diffusion Framework for Human Motion

2026-06-19 · Dongseok Shim, Julian Tanke, Kengo Uchida, Christian Simon 외 arxiv

Human motion generation has been widely studied across diverse input modalities, text, music, and video, and recent efforts have unified these into single multimodal frameworks. However, while morphological factors such …

TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models

2025-12-01 · Zhiheng Liu, Weiming Ren, Haozhe Liu, Zijian Zhou 외 arxiv

Unified multimodal models (UMMs) aim to jointly perform multimodal understanding and generation within a single framework. We present TUNA, a native UMM that builds a unified continuous visual representation by cascading…

Video GenerationImage Editing

Does Unification Come at a Cost? Uni-SafeBench: A Safety Benchmark for Unified Multimodal Large Models

2026-04-01 · Zixiang Peng, Yongxiu Xu, Qinyi Zhang, Jiexun Shen 외 arxiv

Unified Multimodal Large Models (UMLMs) integrate understanding and generation capabilities within a single architecture. While unified architectures expand multimodal capabilities, their safety implications remain impor…

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

2026-08-03 · Junliang Ye, Kenkun Liu, Guocun Wang, Yang Li 외 hf

Recent advances in image generation have demonstrated the potential of unified multimodal models that integrate understanding, generation, and editing. However, unified 3D modeling remains constrained by scarce multimoda…

Image Generation3D Generation

EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning

2025-11-03 · Xinyan Cai, Shiguang Wu, Dafeng Chi, Yuzheng Zhuang 외 arxiv

In complex embodied long-horizon manipulation tasks, effective task decomposition and execution require synergistic integration of textual logical reasoning and visual-spatial imagination to ensure efficient and accurate…

multimodal generationLogical Reasoning