paper-with-me

Papers

Setting the Stage: Text-Driven Scene-Consistent Image Generation

2025-12-14 · Cong Xie, Che Wang, Yan Zhang, Ruiqi Yu, Han Zou, Zheng Pan, Zhenpeng Zhan arxiv

We focus on the foundational task of Scene Staging: given a reference scene image and a text condition specifying an actor category to be generated in the scene and its spatial relation to the scene, the goal is to synthesize an output image that preserves the same scene identity as the reference image while correctly generating the actor according to the spatial relation described in the text. Existing methods struggle with this task, largely due to the scarcity of high-quality paired data and unconstrained generation objectives. To overcome the data bottleneck, we propose a novel data construction pipeline that combines real-world photographs, entity removal, and image-to-video diffusion models to generate training pairs with diverse scenes, viewpoints and correct entity-scene relationships. Furthermore, we introduce a novel correspondence-guided attention loss that leverages cross-view cues to enforce spatial alignment with the reference scene. Experiments on our scene-consistent benchmark show that our approach achieves better scene alignment and text-image alignment than state-of-the-art baselines, according to both automatic metrics and human preference studies. Our method generates images with diverse viewpoints and compositions while faithfully following the textual instructions and preserving the reference scene identity.

📄 PDF Abstract BibTeX arXiv:2512.12598

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting

2024-08-25 · Wenrui Li, Fucheng Cai, Yapeng Mi, Zhe Yang 외

Text-driven 3D scene generation has seen significant advancements recently. However, most existing methods generate single-view images using generative models and then stitch them together in 3D space. This independent g…

3DGSImage GenerationScene Generation

UNITS: Unsupervised Intermediate Training Stage for Scene Text Detection

2022-05-10 · Youhui Guo, Yu Zhou, Xugong Qin, Enze Xie 외

Recent scene text detection methods are almost based on deep learning and data-driven. Synthetic data is commonly adopted for pre-training due to expensive annotation cost. However, there are obvious domain discrepancies…

Scene Text DetectionText Detection

PSGS: Text-driven Panorama Sliding Scene Generation via Gaussian Splatting

2026-01-31 · Xin Zhang, Shen Chen, Jiale Zhou, Lei Li arxiv

Generating realistic 3D scenes from text is crucial for immersive applications like VR, AR, and gaming. While text-driven approaches promise efficiency, existing methods suffer from limited 3D-text data and inconsistent …

Scene GenerationPoint Clouds

Generative Photography: Scene-Consistent Camera Control for Realistic Text-to-Image Synthesis

2024-12-03 · CVPR 2025 1 · Yu Yuan, Xijun Wang, Yichen Sheng, Prateek Chennuri 외

Image generation today can produce somewhat realistic images from text prompts. However, if one asks the generator to synthesize a particular camera setting such as creating different fields of view using a 24mm lens ver…

Image Generation

Drag4D: Align Your Motion with Text-Driven 3D Scene Generation

2025-09-26 · Minjun Kang, Inkyu Shin, Taeyeop Lee, In So Kweon 외 arxiv

We introduce Drag4D, an interactive framework that integrates object motion control within text-driven 3D scene generation. This framework enables users to define 3D trajectories for the 3D objects generated from a singl…

Scene Generation