paper-with-me

홈 › Papers

A$^\text{T}$A: Adaptive Transformation Agent for Text-Guided Subject-Position Variable Background Inpainting

2025-04-02 · Yizhe Tang, Zhimin Sun, Yuzhen Du, Ran Yi, Guangben Lu, Teng Hu, Luying Li, Lizhuang Ma, Fangyuan Zou

Image inpainting aims to fill the missing region of an image. Recently, there has been a surge of interest in foreground-conditioned background inpainting, a sub-task that fills the background of an image while the foreground subject and associated text prompt are provided. Existing background inpainting methods typically strictly preserve the subject's original position from the source image, resulting in inconsistencies between the subject and the generated background. To address this challenge, we propose a new task, the "Text-Guided Subject-Position Variable Background Inpainting", which aims to dynamically adjust the subject position to achieve a harmonious relationship between the subject and the inpainted background, and propose the Adaptive Transformation Agent (A$^\text{T}$A) for this task. Firstly, we design a PosAgent Block that adaptively predicts an appropriate displacement based on given features to achieve variable subject-position. Secondly, we design the Reverse Displacement Transform (RDT) module, which arranges multiple PosAgent blocks in a reverse structure, to transform hierarchical feature maps from deep to shallow based on semantic information. Thirdly, we equip A$^\text{T}$A with a Position Switch Embedding to control whether the subject's position in the generated image is adaptively predicted or fixed. Extensive comparative experiments validate the effectiveness of our A$^\text{T}$A approach, which not only demonstrates superior inpainting capabilities in subject-position variable inpainting, but also ensures good performance on subject-position fixed inpainting.

📄 PDF Abstract BibTeX arXiv:2504.01603

Code (0)

등록된 구현이 없습니다.

Tasks

Image InpaintingPosition

Methods 이 논문이 사용한 방법론

Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.

Similar Papers 제목 키워드 기반

ATA: Adaptive Transformation Agent for Text-Guided Subject-Position Variable Background Inpainting

2025-01-01 · CVPR 2025 1 · Yizhe Tang, Zhimin Sun, Yuzhen Du, Ran Yi 외

Image inpainting aims to fill the missing region of an image.Recently, there has been a surge of interest in foreground-conditioned background inpainting, a sub-task that fills the background of an image while the fo…

Image InpaintingPosition

Text-Guided Unsupervised Latent Transformation for Multi-Attribute Image Manipulation

2023-01-01 · CVPR 2023 1 · Xiwen Wei, Zhen Xu, Cheng Liu, Si Wu 외

Great progress has been made in StyleGAN-based image editing. To associate with preset attributes, most existing approaches focus on supervised learning for semantically meaningful latent space traversal directions, …

AttributeImage ManipulationSemantic SimilaritySemantic Textual Similarity

RLinf: Flexible and Efficient Large-scale Reinforcement Learning via Macro-to-Micro Flow Transformation

2025-09-19 · Chao Yu, Yuanqing Wang, Zhen Guo, Hao Lin 외 arxiv

Reinforcement learning (RL) has demonstrated immense potential in advancing artificial general intelligence, agentic intelligence, and embodied intelligence. However, the inherent heterogeneity and dynamicity of RL workf…

Reinforcement Learning

Situational Perception Guided Image Matting

2022-04-20 · Bo Xu, Jiake Xie, Han Huang, Ziwen Li 외

Most automatic matting methods try to separate the salient foreground from the background. However, the insufficient quantity and subjective bias of the current existing matting datasets make it difficult to fully explor…

Image MattingObject

Content-Adaptive Image Retouching Guided by Attribute-Based Text Representation

2025-12-10 · Hancheng Zhu, Xinyu Liu, Rui Yao, Kunyang Sun 외 arxiv

Image retouching has received significant attention due to its ability to achieve high-quality visual content. Existing approaches mainly rely on uniform pixel-wise color mapping across entire images, neglecting the inhe…