paper-with-me

Papers

MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration

2024-08-20 · Yanbo Ding, Shaobin Zhuang, Kunchang Li, Zhengrong Yue, Yu Qiao, Yali Wang

Despite recent advancements in text-to-image generation, most existing methods struggle to create images with multiple objects and complex spatial relationships in the 3D world. To tackle this limitation, we introduce a generic AI system, namely MUSES, for 3D-controllable image generation from user queries. Specifically, our MUSES addresses this challenging task by developing a progressive workflow with three key components, including (1) Layout Manager for 2D-to-3D layout lifting, (2) Model Engineer for 3D object acquisition and calibration, (3) Image Artist for 3D-to-2D image rendering. By mimicking the collaboration of human professionals, this multi-modal agent pipeline facilitates the effective and automatic creation of images with 3D-controllable objects, through an explainable integration of top-down planning and bottom-up generation. Additionally, we find that existing benchmarks lack detailed descriptions of complex 3D spatial relationships of multiple objects. To fill this gap, we further construct a new benchmark of T2I-3DisBench (3D image scene), which describes diverse 3D image scenes with 50 detailed prompts. Extensive experiments show the state-of-the-art performance of MUSES on both T2I-CompBench and T2I-3DisBench, outperforming recent strong competitors such as DALL-E 3 and Stable Diffusion 3. These results demonstrate a significant step of MUSES forward in bridging natural language, 2D image generation, and 3D world. Our codes are available at the following link: https://github.com/DINGYANB/MUSES.

📄 PDF Abstract BibTeX arXiv:2408.10605

Code (1)

DINGYANB/MUSES 공식 구현 pytorch

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MUSES: The Multi-Sensor Semantic Perception Dataset for Driving under Uncertainty

2024-01-23 · Tim Brödermann, David Bruggemann, Christos Sakaridis, Kevin Ta 외

Achieving level-5 driving automation in autonomous vehicles necessitates a robust semantic visual perception system capable of parsing data from different sensors across diverse conditions. However, existing semantic per…

Autonomous VehiclesObject DetectionPanoptic SegmentationSemantic Segmentation+1

Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without Training

2026-01-06 · Hexiao Lu, Xiaokun Sun, Zeyu Cai, Hao Guo 외 arxiv

We present Muses, the first training-free method for fantastic 3D creature generation in a feed-forward paradigm. Previous methods, which rely on part-aware optimization, manual assembly, or 2D image generation, often pr…

3D Object EditingImage Generation

Multi-shot Temporal Event Localization: a Benchmark

2020-12-17 · CVPR 2021 1 · Xiaolong Liu, Yao Hu, Song Bai, Fei Ding 외

Current developments in temporal event or action localization usually target actions captured by a single camera. However, extensive events or actions in the wild may be captured as a sequence of shots by multiple camera…

Action LocalizationTemporal Action Localization

FlexEControl: Flexible and Efficient Multimodal Control for Text-to-Image Generation

2024-05-08 · Xuehai He, Jian Zheng, Jacob Zhiyuan Fang, Robinson Piramuthu 외

Controllable text-to-image (T2I) diffusion models generate images conditioned on both text prompts and semantic inputs of other modalities like edge maps. Nevertheless, current controllable T2I methods commonly face chal…

Image GenerationText to Image GenerationText-to-Image Generation

Stance-Driven Multimodal Controlled Statement Generation: New Dataset and Task

2025-04-04 · Bingqian Wang, Quan Fang, Jiachen Sun, Xiaoxiao Ma

Formulating statements that support diverse or controversial stances on specific topics is vital for platforms that enable user expression, reshape political discourse, and drive social critique and information dissemina…

Marketingmultimodal generationStance DetectionText Generation