paper-with-me

Papers

XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation

2025-06-26 · Bowen Chen, Mengyi Zhao, Haomiao Sun, Li Chen, Xu Wang, Kang Du, Xinglong Wu

Achieving fine-grained control over subject identity and semantic attributes (pose, style, lighting) in text-to-image generation, particularly for multiple subjects, often undermines the editability and coherence of Diffusion Transformers (DiTs). Many approaches introduce artifacts or suffer from attribute entanglement. To overcome these challenges, we propose a novel multi-subject controlled generation model XVerse. By transforming reference images into offsets for token-specific text-stream modulation, XVerse allows for precise and independent control for specific subject without disrupting image latents or features. Consequently, XVerse offers high-fidelity, editable multi-subject image synthesis with robust control over individual subject characteristics and semantic attributes. This advancement significantly improves personalized and complex scene generation capabilities.

📄 PDF Abstract BibTeX arXiv:2506.21416

Code (1)

bytedance/xverse 공식 구현 pytorch

Tasks

AttributeImage GenerationScene GenerationText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

CRAFT: Constrained Reward via Attention Fine-Tuning for Subject Personalization without Composed Targets

2026-08-14 · Jihun Park, Kyoungmin Lee, Jongmin Gim, Hyeonseo Jo 외 arxiv

Subject-driven image personalization---generating new images that preserve the identity of one or several reference subjects in novel scenes---is a foundational capability for modern visual content creation. It is curren…

CopyCat: Improving Fine-Grained Subject Consistency in Subject-to-Image Models within Seconds

2026-08-01 · Peng Zheng, Ruiqi Liu, Rui Ma, Zuxuan Wu arxiv

Recent subject-to-image models have achieved impressive progress in personalized image generation, yet they still struggle to preserve fine-grained subject-specific details. A major reason is the lack of high-quality fin…

Personalized Image Generation

When Identities Collapse: A Stress-Test Benchmark for Multi-Subject Personalization

2026-03-27 · Zhihan Chen, Yuhuan Zhao, Yijie Zhu, Xinyu Yao arxiv

Subject-driven text-to-image diffusion models have achieved remarkable success in preserving single identities, yet their ability to compose multiple interacting subjects remains largely unexplored and highly challenging…

DexVerse: A Modular Benchmark for Multi-Task, Multi-Embodiment Dexterous Manipulation

2026-07-09 · Yunchao Yao, Zhuxiu Xu, Tianqi Zhang, Zixian Liu 외 arxiv

Building general-purpose dexterous manipulation policies requires benchmarks that go beyond isolated tasks to systematically evaluate policies across diverse interaction modes, sensory conditions, and robot embodiments. …

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement

2025-06-09 · Teng Hu, Zhentao Yu, Zhengguang Zhou, Jiangning Zhang 외

Despite recent advances in video generation, existing models still lack fine-grained controllability, especially for multi-subject customization with consistent identity and interaction. In this paper, we propose PolyViv…

Video Generation