paper-with-me

Papers

Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without Training

2026-01-06 · Hexiao Lu, Xiaokun Sun, Zeyu Cai, Hao Guo, Ying Tai, Jian Yang, Zhenyu Zhang arxiv

We present Muses, the first training-free method for fantastic 3D creature generation in a feed-forward paradigm. Previous methods, which rely on part-aware optimization, manual assembly, or 2D image generation, often produce unrealistic or incoherent 3D assets due to the challenges of intricate part-level manipulation and limited out-of-domain generation. In contrast, Muses leverages the 3D skeleton, a fundamental representation of biological forms, to explicitly and rationally compose diverse elements. This skeletal foundation formalizes 3D content creation as a structure-aware pipeline of design, composition, and generation. Muses begins by constructing a creatively composed 3D skeleton with coherent layout and scale through graph-constrained reasoning. This skeleton then guides a voxel-based assembly process within a structured latent space, integrating regions from different objects. Finally, image-guided appearance modeling under skeletal conditions is applied to generate a style-consistent and harmonious texture for the assembled shape. Extensive experiments establish Muses' state-of-the-art performance in terms of visual fidelity and alignment with textual descriptions, and potential on flexible 3D object editing. Project page: https://luhexiao.github.io/Muses.github.io/.

📄 PDF Abstract BibTeX arXiv:2601.03256

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object EditingImage Generation

Similar Papers 제목 키워드 기반

Using Small MUSes to Explain How to Solve Pen and Paper Puzzles

2021-04-30 · Joan Espasa, Ian P. Gent, Ruth Hoffmann, Christopher Jefferson 외

In this paper, we present Demystify, a general tool for creating human-interpretable step-by-step explanations of how to solve a wide range of pen and paper puzzles from a high-level logical description. Demystify is bas…

Multi-shot Temporal Event Localization: a Benchmark

2020-12-17 · CVPR 2021 1 · Xiaolong Liu, Yao Hu, Song Bai, Fei Ding 외

Current developments in temporal event or action localization usually target actions captured by a single camera. However, extensive events or actions in the wild may be captured as a sequence of shots by multiple camera…

Action LocalizationTemporal Action Localization

MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration

2024-08-20 · Yanbo Ding, Shaobin Zhuang, Kunchang Li, Zhengrong Yue 외

Despite recent advancements in text-to-image generation, most existing methods struggle to create images with multiple objects and complex spatial relationships in the 3D world. To tackle this limitation, we introduce a …

Image GenerationText to Image GenerationText-to-Image Generation

Emerging-properties Mapping Using Spatial Embedding Statistics: EMUSES

2024-06-20 · Chris Foulon, Marcela Ovando-Tellez, Lia Talozzi, Maurizio Corbetta 외

Understanding complex phenomena often requires analyzing high-dimensional data to uncover emergent properties that arise from multifactorial interactions. Here, we present EMUSES (Emerging-properties Mapping Using Spatia…

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers

2025-07-17 · Qiang Wang, Mengchao Wang, Fan Jiang, Yaqi Fan 외 arxiv

Producing expressive facial animations from static images is a challenging task. Prior methods relying on explicit geometric priors (e.g., facial landmarks or 3DMM) often suffer from artifacts in cross reenactment and st…