paper-with-me

홈 › Papers

COMPASS: Grounding Composition-Intent Guidance in Unified Multimodal Models

2026-06-27 · Ziqi Zhou, Weize Quan, Mining Tan, Zhihan Chen, Dandan Zheng, Jingdong Chen, Jun Zhou, Weiming Dong, Dong-Ming Yan arxiv

Composition is a high-level visual intent that governs where subjects are placed and how a scene is organized, yet current unified multimodal models remain unreliable at fine-grained composition recognition and struggle to turn such intent into controllable generation. We present COMPASS, the first unified multimodal framework that grounds composition-intent control in a single system spanning both composition perception and composition-guided generation, with a shared expert token $τ_c$ as the central intent anchor. On the perception side, COMPASS injects composition expertise into an MoE backbone in a minimally invasive manner and distills the inferred intent into $τ_c$. On the generation side, COMPASS reuses $τ_c$ as a global conditioning signal that steers the denoising trajectory, effectively converting passive composition analysis into explicit layout control. To support systematic instruction-following composition learning and evaluation at scale, we construct Comp-11, a large-scale dataset with an 11-class taxonomy and reasoning-augmented annotations. Extensive experiments show that COMPASS substantially improves category-level composition understanding and delivers more composition-consistent, prompt-faithful generation than strong baselines.

📄 PDF Abstract BibTeX arXiv:2606.28696

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CompassAD: Intent-Driven 3D Affordance Grounding in Functionally Competing Objects

2026-04-02 · Jingliang Li, Jindou Jia, Tuo An, Chuhao Zhou 외 arxiv

When told to "cut the cake," a robot must choose the knife over nearby scissors, despite both objects affording the same cutting function. In real-world scenes, multiple objects may share identical affordances, yet only …

Visual Intention Grounding for Egocentric Assistants

2025-04-18 · Pengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li 외

Visual grounding associates textual descriptions with objects in an image. Conventional methods target third-person image inputs and named object queries. In applications such as AI assistants, the perspective shifts -- …

ObjectVisual Grounding

OmniWeaving: Towards Unified Video Generation with Free-form Composition and Reasoning

2026-03-25 · Kaihang Pan, Qi Tian, Jianwei Zhang, Weijie Kong 외 arxiv

While proprietary systems such as Seedance-2.0 have achieved remarkable success in omni-capable video generation, open-source alternatives significantly lag behind. Most academic models remain heavily fragmented, and the…

Video Generation

BrickCraft: Visuomotor Skill Composition with Situated Manual Guidance for Long-Horizon Interlocking Brick Assembly

2026-05-08 · Jichuan Yu, Bowei Li, Zhenran Tang, Guanxing Lu 외 arxiv

Autonomous robotic assembly of interlocking bricks demands seamless integration of long-horizon task reasoning, spatial grounding, and fine-grained manipulation. This paper presents BrickCraft, a compositional framework …

TFG: Unified Training-Free Guidance for Diffusion Models

2024-09-24 · Haotian Ye, Haowei Lin, Jiaqi Han, Minkai Xu 외

Given an unconditional diffusion model and a predictor for a target property of interest (e.g., a classifier), the goal of training-free guidance is to generate samples with desirable target properties without additional…