paper-with-me

Papers

Beyond Single Prompts: Synergistic Fusion and Arrangement for VICL

2026-01-15 · Wenwen Liao, Jianbo Yu, Yuansong Wang, Shifu Yan, Xiaofeng Yang arxiv

Vision In-Context Learning (VICL) enables inpainting models to quickly adapt to new visual tasks from only a few prompts. However, existing methods suffer from two key issues: (1) selecting only the most similar prompt discards complementary cues from other high-quality prompts; and (2) failing to exploit the structured information implied by different prompt arrangements. We propose an end-to-end VICL framework to overcome these limitations. Firstly, an adaptive Fusion Module aggregates critical patterns and annotations from multiple prompts to form more precise contextual prompts. Secondly, we introduce arrangement-specific lightweight MLPs to decouple layout priors from the core model, while minimally affecting the overall model. In addition, an bidirectional fine-tuning mechanism swaps the roles of query and prompt, encouraging the model to reconstruct the original prompt from fused context and thus enhancing collaboration between the fusion module and the inpainting model. Experiments on foreground segmentation, single-object detection, and image colorization demonstrate superior results and strong cross-task generalization of our method.

📄 PDF Abstract BibTeX arXiv:2601.10117

Code (0)

등록된 구현이 없습니다.

Tasks

Image ColorizationObject Detection

Similar Papers 제목 키워드 기반

Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion Generation

2024-10-01 · ACMMM24 2024 10 · Bohong Chen, Yumeng Li, Yao-Xiang Ding, Tianjia Shao 외

Current co-speech motion generation approaches usually focus on upper body gestures following speech contents only, while lacking supporting the elaborate control of synergistic full-body motion based on text prompts, su…

Gesture GenerationMotion Generation

SmartSpatial: Enhancing the 3D Spatial Arrangement Capabilities of Stable Diffusion Models and Introducing a Novel 3D Spatial Evaluation Framework

2025-01-01 · Mao Xun Huang, Hen-Hsen Huang

Stable Diffusion models have made remarkable strides in generating photorealistic images from text prompts but often falter when tasked with accurately representing complex spatial arrangements, particularly involving in…

Dependency ParsingImage Generation

SynGR: Unleashing the Potential of Cross-Modal Synergy for Generative Recommendation

2026-05-18 · Wei Chen, Xingyu Guo, Shuang Li, Fuwei Zhang 외 arxiv

Generative Recommendation (GR) has emerged as a promising paradigm by formulating item recommendation as a sequence-to-sequence generation task over item identifiers. Recent studies have incorporated multimodal signals t…

BeyondFacial: Identity-Preserving Personalized Generation Beyond Facial Close-ups

2025-11-15 · Songsong Zhang, Chuanqi Tang, Hongguang Zhang, Guijian Tang 외 arxiv

Identity-Preserving Personalized Generation (IPPG) has advanced film production and artistic creation, yet existing approaches overemphasize facial regions, resulting in outputs dominated by facial close-ups.These method…

SPICE: A Synergistic, Precise, Iterative, and Customizable Image Editing Workflow

2025-04-13 · Kenan Tang, Yanhong Li, Yao Qin

Recent prompt-based image editing models have demonstrated impressive prompt-following capability at structural editing tasks. However, existing models still fail to perform local edits, follow detailed editing prompts, …