paper-with-me

홈 › Papers

ST-LDM: A Universal Framework for Text-Grounded Object Generation in Real Images

2024-03-15 · Xiangtian Xue, Jiasong Wu, Youyong Kong, Lotfi Senhadji, Huazhong Shu

We present a novel image editing scenario termed Text-grounded Object Generation (TOG), defined as generating a new object in the real image spatially conditioned by textual descriptions. Existing diffusion models exhibit limitations of spatial perception in complex real-world scenes, relying on additional modalities to enforce constraints, and TOG imposes heightened challenges on scene comprehension under the weak supervision of linguistic information. We propose a universal framework ST-LDM based on Swin-Transformer, which can be integrated into any latent diffusion model with training-free backward guidance. ST-LDM encompasses a global-perceptual autoencoder with adaptable compression scales and hierarchical visual features, parallel with deformable multimodal transformer to generate region-wise guidance for the subsequent denoising process. We transcend the limitation of traditional attention mechanisms that only focus on existing visual features by introducing deformable feature alignment to hierarchically refine spatial positioning fused with multi-scale visual and linguistic information. Extensive Experiments demonstrate that our model enhances the localization of attention mechanisms while preserving the generative capabilities inherent to diffusion models.

📄 PDF Abstract BibTeX arXiv:2403.10004

Code (0)

등록된 구현이 없습니다.

Tasks

Denoising

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Focus 설명 없음
Latent Diffusion Model Diffusion models applied to latent spaces, which are normally built with (Variational) Autoencoders.

Similar Papers 제목 키워드 기반

UniBiomed: A Universal Foundation Model for Grounded Biomedical Image Interpretation

2025-04-30 · Linshan Wu, Yuxiang Nie, Sunan He, Jiaxin Zhuang 외

Multi-modal interpretation of biomedical images opens up novel opportunities in biomedical image analysis. Conventional AI approaches typically rely on disjointed training, i.e., Large Language Models (LLMs) for clinical…

DiagnosticLarge Language ModelQuestion AnsweringText Generation+1

Text-to-3D Gaussian Splatting with Physics-Grounded Motion Generation

2024-12-07 · Wenqing Wang, Yun Fu

Text-to-3D generation is a valuable technology in virtual reality and digital content creation. While recent works have pushed the boundaries of text-to-3D generation, producing high-fidelity 3D objects with inefficient …

3D GenerationLanguage ModelingLanguage ModellingLarge Language Model+3

GVDIFF: Grounded Text-to-Video Generation with Diffusion Models

2024-07-02 · Huanzhang Dou, Ruixiang Li, Wei Su, Xi Li

In text-to-video (T2V) generation, significant attention has been directed toward its development, yet unifying discrete and continuous grounding conditions in T2V generation remains under-explored. This paper proposes a…

Text-to-Video GenerationVideo Generation

Large Language Models are Universal Reasoners for Visual Generation

2026-05-05 · Sucheng Ren, Chen Chen, Zhenbang Wang, Liangchen Song 외 arxiv

Text-to-image generation has advanced rapidly with diffusion models, progressing from CLIP and T5 conditioning to unified systems where a single LLM backbone handles both visual understanding and generation. Despite the …

Text-to-Image Generation

Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking Avatars

2026-02-02 · Youliang Zhang, Zhengguang Zhou, Zhentao Yu, Ziyao Huang 외 arxiv

Generating talking avatars is a fundamental task in video generation. Although existing methods can generate full-body talking avatars with simple human motion, extending this task to grounded human-object interaction (G…

Video Generation