paper-with-me

홈 › Papers

Text-Image Conditioned 3D Generation

2026-03-22 · Jiazhong Cen, Jiemin Fang, Sikuang Li, Guanjun Wu, Chen Yang, Taoran Yi, Zanwei Zhou, Zhikuan Bao, Lingxi Xie, Wei Shen, Qi Tian arxiv

High-quality 3D assets are essential for VR/AR, industrial design, and entertainment, motivating growing interest in generative models that create 3D content from user prompts. Most existing 3D generators, however, rely on a single conditioning modality: image-conditioned models achieve high visual fidelity by exploiting pixel-aligned cues but suffer from viewpoint bias when the input view is limited or ambiguous, while text-conditioned models provide broad semantic guidance yet lack low-level visual detail. This limits how users can express intent and raises a natural question: can these two modalities be combined for more flexible and faithful 3D generation? Our diagnostic study shows that even simple late fusion of text- and image-conditioned predictions outperforms single-modality models, revealing strong cross-modal complementarity. We therefore formalize Text-Image Conditioned 3D Generation, which requires joint reasoning over a visual exemplar and a textual specification. To address this task, we introduce TIGON, a minimalist dual-branch baseline with separate image- and text-conditioned backbones and lightweight cross-modal fusion. Extensive experiments show that text-image conditioning consistently improves over single-modality methods, highlighting complementary vision-language guidance as a promising direction for future 3D generation research. Project page: https://jumpat.github.io/tigon-page

📄 PDF Abstract BibTeX arXiv:2603.21295

Code (0)

등록된 구현이 없습니다.

Tasks

3D Generation

Similar Papers 제목 키워드 기반

TextAtlas5M: A Large-scale Dataset for Dense Text Image Generation

2025-02-11 · Alex Jinpeng Wang, Dongxing Mao, Jiawei Zhang, Weiming Han 외

Text-conditioned image generation has gained significant attention in recent years and are processing increasingly longer and comprehensive text prompt. In everyday life, dense and intricate text appears in contexts like…

Image Generation

AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation

2023-05-22 · Interspeech 2023 5 · Guy Yariv, Itai Gat, Lior Wolf, Yossi Adi 외

In recent years, image generation has shown a great leap in performance, where diffusion models play a central role. Although generating high-quality images, such models are mainly conditioned on textual descriptions. Th…

audio-visual learningImage GenerationText to Image GenerationText-to-Image Generation

Semantically Invariant Text-to-Image Generation

2018-09-27 · Shagan Sah, Dheeraj Peri, Ameya Shringi, Chi Zhang 외

Image captioning has demonstrated models that are capable of generating plausible text given input images or videos. Further, recent work in image generation has shown significant improvements in image quality when text …

Image CaptioningImage GenerationText to Image GenerationText-to-Image Generation

XGPT: Cross-modal Generative Pre-Training for Image Captioning

2020-03-03 · Qiaolin Xia, Haoyang Huang, Nan Duan, Dong-dong Zhang 외

While many BERT-based cross-modal pre-trained models produce excellent results on downstream understanding tasks like image-text retrieval and VQA, they cannot be applied to generation tasks directly. In this paper, we p…

Data AugmentationDenoisingImage CaptioningImage Retrieval+7

UV-IDM: Identity-Conditioned Latent Diffusion Model for Face UV-Texture Generation

2024-01-01 · CVPR 2024 1 · Hong Li, Yutang Feng, Song Xue, Xuhui Liu 외

3D face reconstruction aims at generating high-fidelity 3D face shapes and textures from single-view or multi-view images. However current prevailing facial texture generation methods generally suffer from low-qualit…

3D Face ReconstructionFace ModelFace ReconstructionTexture Synthesis