paper-with-me

Papers

IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation

2024-09-12 · Yinwei Wu, Xianpan Zhou, Bing Ma, Xuefeng Su, Kai Ma, Xinchao Wang

While Text-to-Image (T2I) diffusion models excel at generating visually appealing images of individual instances, they struggle to accurately position and control the features generation of multiple instances. The Layout-to-Image (L2I) task was introduced to address the positioning challenges by incorporating bounding boxes as spatial control signals, but it still falls short in generating precise instance features. In response, we propose the Instance Feature Generation (IFG) task, which aims to ensure both positional accuracy and feature fidelity in generated instances. To address the IFG task, we introduce the Instance Feature Adapter (IFAdapter). The IFAdapter enhances feature depiction by incorporating additional appearance tokens and utilizing an Instance Semantic Map to align instance-level features with spatial locations. The IFAdapter guides the diffusion process as a plug-and-play module, making it adaptable to various community models. For evaluation, we contribute an IFG benchmark and develop a verification pipeline to objectively compare models' abilities to generate instances with accurate positioning and features. Experimental results demonstrate that IFAdapter outperforms other models in both quantitative and qualitative evaluations.

📄 PDF Abstract BibTeX arXiv:2409.08240

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationText to Image GenerationText-to-Image Generation

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…
Adapter 설명 없음
ALIGN In the ALIGN method, visual and language representations are jointly trained from noisy image alt-text data. The image and text encoders are learned via contrastive loss…

Similar Papers 제목 키워드 기반

ConFusion: Continuous Fusion Space Learning for Fine-Grained Controllable Infrared and Visible Image Fusion

2026-07-26 · Guo Yurong, He Yufei, Li Yonghao, Chang Dongliang 외 arxiv

Controllable infrared-visible image fusion aims to integrate complementary thermal and structural information with flexible region-aware modulation, producing fused images that adapt to diverse user requirements and down…

OmniACBench: A Benchmark for Evaluating Context-Grounded Acoustic Control in Omni-Modal Models

2026-03-25 · Seunghee Kim, Bumkyu Park, Kyudan Jung, Joosung Lee 외 arxiv

Most testbeds for omni-modal models assess multimodal understanding via textual outputs, leaving it unclear whether these models can properly speak their answers. To study this, we introduce OmniACBench, a benchmark for …

Re$^2$Math: Benchmarking Theorem Retrieval in Research-Level Mathematics

2026-05-09 · Zicheng Lyu, Wenjie Yang, Shengzhong Zhang, Zengfeng Huang arxiv

Large language models are increasingly capable at closed-world mathematical reasoning, but research assistance also requires source-grounded use of the literature. When a proof reaches a non-trivial step, a useful assist…

Mathematical Reasoning

ConsistCompose: Unified Multimodal Layout Control for Image Composition

2025-11-23 · Xuanke Shi, Boxuan Li, Xiaoyang Han, Zhongang Cai 외 arxiv

Unified multimodal models that couple visual understanding with image generation have advanced rapidly, yet most systems still focus on visual grounding-aligning language with image regions-while their generative counter…

Visual GroundingImage Generation

IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot Manner

2025-01-01 · CVPR 2025 1 · YuYang Huang, Yabo Chen, Li Ding, Xiaopeng Zhang 외

Controllability of video generation has been recently concerned in addition to the quality of generated videos. The main challenge to controllable video generation is to synthesize videos based on user-specified inst…

Motion GenerationText-to-Video GenerationVideo Generation