paper-with-me

홈 › Papers

Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes

2026-04-16 · Victoria Yue Chen, Emery Pierson, Léopold Maillard, Maks Ovsjanikov arxiv

Text-driven inversion of generative models is a core paradigm for manipulating 2D or 3D content, unlocking numerous applications such as text-based editing, style transfer, or inverse problems. However, it relies on the assumption that generative models remain sensitive to natural language prompts. We demonstrate that for state-of-the-art native text-to-3D generative models, this assumption often collapses. We identify a critical failure mode where generation trajectories are drawn into latent "sink traps": regions where the model becomes insensitive to prompt modifications. In these regimes, changes to the input text fail to alter internal representations in a way that alters the output geometry. Crucially, we observe that this is not a limitation of the model's \textit{geometric} expressivity; the same generative models possess the ability to produce a vast diversity of shapes but, as we demonstrate, become insensitive to out-of-distribution \textit{text} guidance. We investigate this behavior by analyzing the sampling trajectories of the generative model, and find that complex geometries can still be represented and produced by leveraging the model's unconditional generative prior. This leads to a more robust framework for text-based 3D shape editing that bypasses latent sinks by decoupling a model's geometric representation power from its linguistic sensitivity. Our approach addresses the limitations of current 3D pipelines and enables high-fidelity semantic manipulation of out-of-distribution 3D shapes. Project webpage: https://daidedou.sorpi.fr/publication/beyondprompts

📄 PDF Abstract BibTeX arXiv:2604.14914

Code (0)

등록된 구현이 없습니다.

Tasks

Style Transfer

Similar Papers 제목 키워드 기반

When Does High-CFG Diffusion Inversion Fail? A Controlled Study of Prompt--Latent Interactions

2026-07-06 · Yan Zeng, Yusuke Hosoya, Huyen T. T. Tran, Takayuki Okatani arxiv

Text-guided diffusion inversion is central to image editing, where an image is mapped to an initial latent and then edited by replaying the denoising process under a modified prompt. In practice, however, inversion is of…

Image Editing

Video-P2P: Video Editing with Cross-attention Control

2023-03-08 · CVPR 2024 1 · Shaoteng Liu, Yuechen Zhang, Wenbo Li, Zhe Lin 외

This paper presents Video-P2P, a novel framework for real-world video editing with cross-attention control. While attention control has proven effective for image editing with pre-trained image generation models, there a…

Image GenerationVideo EditingVideo Generation

Video-P2P: Video Editing with Cross-attention Control

2023-03-08 · arXiv 2023 3 · Shaoteng Liu;Yuechen Zhang;Wenbo Li;Zhe Lin;Jiaya Jia

This paper presents Video-P2P, a novel framework for real-world video editing with cross-attention control. While attention control has proven effective for image editing with pre-trained image generation models, there a…

Image GenerationVideo EditingVideo Generation

Joint 3D Gravity and Magnetic Inversion via Rectified Flow and Ginzburg-Landau Guidance

2026-03-06 · Dhruman Gupta, Yashas Shende, Aritra Das, Chanda Grover Kamra 외 arxiv

Subsurface ore detection is of paramount importance given the rising depletion of shallow mineral resources in recent years. It is crucial to explore approaches that go beyond the limitations of traditional geological ex…

Dual Inversion for Text-to-Image Diffusion Models: From Both Prompt and Noise Perspectives

2026-07-29 · Xiaolong Liu, Junjian Li, Yuan Xiao, Jiaqi Deng 외 arxiv

Prompt inversion, as a typical reverse engineering technique, enables text-to-image (T2I) diffusion models to generate the desired target images without extensive prompt engineering. However, existing prompt inversion me…

Prompt EngineeringImage Editing