paper-with-me

홈 › Papers

FOCUS: Optimal Control for Multi-Entity World Modeling in Text-to-Image Generation

2025-10-02 · Eric Tillmann Bill, Enis Simsar, Thomas Hofmann arxiv

Text-to-image (T2I) models excel on single-entity prompts but struggle with multi-entity scenes, often exhibiting attribute leakage, identity entanglement, and subject omissions. We present a principled theoretical framework that steers sampling toward multi-subject fidelity by casting flow matching (FM) as stochastic optimal control (SOC), yielding a single hyperparameter controlled trade-off between fidelity and object-centric state separation / binding consistency. Within this framework, we derive two architecture-agnostic algorithms: (i) a training-free test-time controller that perturbs the base velocity with a single-pass update, and (ii) Adjoint Matching, a lightweight fine-tuning rule that regresses a control network to a backward adjoint signal. The same formulation unifies prior attention heuristics, extends to diffusion models via a flow--diffusion correspondence, and provides the first fine-tuning route explicitly designed for multi-subject fidelity. In addition, we also introduce FOCUS (Flow Optimal Control for Unentangled Subjects), a probabilistic attention-binding objective compatible with both algorithms. Empirically, on Stable Diffusion 3.5 and FLUX.1, both algorithms consistently improve multi-subject alignment while maintaining base-model style; test-time control runs efficiently on commodity GPUs, and fine-tuned models generalize to unseen prompts.

📄 PDF Abstract BibTeX arXiv:2510.02315

Code (0)

등록된 구현이 없습니다.

Tasks

Text-to-Image Generation

Similar Papers 제목 키워드 기반

Entity Resolution via Batched Oracle Queries

2026-06-23 · Lorenzo Balzotti, Donatella Firmani, Luca Gagliardelli, Giovanni Simonini arxiv

We consider an oracle that processes a limited batch of records at a time and clusters those that refer to the same real-world entity. We study how to interrogate such an oracle to resolve entities in a dataset whose siz…

Entity Resolution

ID-Sim: An Identity-Focused Similarity Metric

2026-04-06 · Julia Chae, Nicholas Kolkin, Jui-Hsien Wang, Richard Zhang 외 arxiv

Humans have remarkable selective sensitivity to identities -- easily distinguishing between highly similar identities, even across significantly different contexts such as diverse viewpoints or lighting. Vision models ha…

Personalized Image Generation

Leveraging Intra-modal and Inter-modal Interaction for Multi-Modal Entity Alignment

2024-04-19 · Zhiwei Hu, Víctor Gutiérrez-Basulto, Zhiliang Xiang, Ru Li 외

Multi-modal entity alignment (MMEA) aims to identify equivalent entity pairs across different multi-modal knowledge graphs (MMKGs). Existing approaches focus on how to better encode and aggregate information from differe…

Contrastive LearningEntity AlignmentKnowledge GraphsMulti-modal Entity Alignment

DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning

2026-03-12 · Yujie Wei, Xinyu Liu, Shiwei Zhang, Hangjie Yuan 외 arxiv

While large-scale diffusion models have revolutionized video synthesis, achieving precise control over both multi-subject identity and multi-granularity motion remains a significant challenge. Recent attempts to bridge t…

Reinforcement Learning

Controlla: Learning Controllability via Graph-Constrained Latent Geometry

2026-05-15 · Jamuna S. Murthy, Amin Karimi Monsefi, Rajiv Ramnath arxiv

Controllable multimodal generation is commonly formulated as an inference-time conditioning problem using prompts, guidance, or auxiliary modules. While effective, such approaches do not explicitly structure how semantic…

multimodal generation