paper-with-me

Papers

ContextGen: Contextual Layout Anchoring for Identity-Consistent Multi-Instance Generation

2025-10-13 · Ruihang Xu, Dewei Zhou, Fan Ma, Yi Yang arxiv

Multi-instance image generation (MIG) remains a significant challenge for modern diffusion models due to key limitations in achieving precise control over object layout and preserving the identity of multiple distinct subjects. To address these limitations, we introduce ContextGen, a novel Diffusion Transformer framework for multi-instance generation that is guided by both layout and reference images. Our approach integrates two key technical contributions: a Contextual Layout Anchoring (CLA) mechanism that incorporates the composite layout image into the generation context to robustly anchor the objects in their desired positions, and Identity Consistency Attention (ICA), an innovative attention mechanism that leverages contextual reference images to ensure the identity consistency of multiple instances. To address the absence of a large-scale, high-quality dataset for this task, we introduce IMIG-100K, the first dataset to provide detailed layout and identity annotations specifically designed for Multi-Instance Generation. Extensive experiments demonstrate that ContextGen sets a new state-of-the-art, outperforming existing methods especially in layout control and identity fidelity.

📄 PDF Abstract BibTeX arXiv:2510.11000

Code (0)

등록된 구현이 없습니다.

Tasks

Image Generation

Similar Papers 제목 키워드 기반

Envisioning Beyond the Few: Disentangled Semantics and Primitives for Few-Shot Atypical Layout-to-Image Generation

2026-05-29 · Nan Bao, Yifan Zhao, Wenzhuang Wang, Jia Li arxiv

The layout-to-image (L2I) task enables fine-grained control over image generation via object categories and spatial layouts. However, existing L2I methods yield fragmented and distorted generations under few-shot atypica…

Layout-to-Image Generation

Lookahead Anchoring: Preserving Character Identity in Audio-Driven Human Animation

2025-10-27 · Junyoung Seo, Rodrigo Mira, Alexandros Haliassos, Stella Bounareli 외 arxiv

Audio-driven human animation models often suffer from identity drift during temporal autoregressive generation, where characters gradually lose their identity over time. One solution is to generate keyframes as intermedi…

ISAP-3D: Identity-Slot Aligned Part-Aware 3D Generation

2026-06-10 · Junlin Hao, Haoshuai Fu, Xibin Song, Wei Li 외 arxiv

Part-aware 3D generation aims to synthesize structured objects with semantically meaningful components, yet often suffers from structural ambiguity due to identity-layout entanglement. Existing methods either infer part …

3D Generation

SAD-GS: Learning Reliable 3D Semantic Gaussian Fields via Dynamic Geo-Semantic Anchoring

2026-06-28 · Yufei Zhang, Chenlu Zhan, Gaoang Wang, Hongwei Wang arxiv

Open-vocabulary 3D semantic Gaussian field learning relies on multi-view 2D supervision, whose semantic targets and spatial assignments are often unreliable. Across varying viewpoints, view-dependent features cause seman…

Semantic Segmentation

Physics-Informed Structure Anchoring With Capture-Aware Prototype Calibration for Cross-Environment RF Fingerprinting

2026-07-06 · Fengchong Yao, Jianbing Li, Qing Liu, Qikun Liu 외 arxiv

Radio frequency fingerprint identification (RFFI) exploits transmitter-specific hardware imperfections as physicallayer identity cues for Internet of Things (IoT) devices, but deep models often degrade across acquisition…

Representation Learning