paper-with-me

Papers

Anywhere: A Multi-Agent Framework for User-Guided, Reliable, and Diverse Foreground-Conditioned Image Generation

2024-04-29 · Tianyidan Xie, Rui Ma, Qian Wang, Xiaoqian Ye, Feixuan Liu, Ying Tai, Zhenyu Zhang, Lanjun Wang, Zili Yi

Recent advancements in image-conditioned image generation have demonstrated substantial progress. However, foreground-conditioned image generation remains underexplored, encountering challenges such as compromised object integrity, foreground-background inconsistencies, limited diversity, and reduced control flexibility. These challenges arise from current end-to-end inpainting models, which suffer from inaccurate training masks, limited foreground semantic understanding, data distribution biases, and inherent interference between visual and textual prompts. To overcome these limitations, we present Anywhere, a multi-agent framework that departs from the traditional end-to-end approach. In this framework, each agent is specialized in a distinct aspect, such as foreground understanding, diversity enhancement, object integrity protection, and textual prompt consistency. Our framework is further enhanced with the ability to incorporate optional user textual inputs, perform automated quality assessments, and initiate re-generation as needed. Comprehensive experiments demonstrate that this modular design effectively overcomes the limitations of existing end-to-end models, resulting in higher fidelity, quality, diversity and controllability in foreground-conditioned image generation. Additionally, the Anywhere framework is extensible, allowing it to benefit from future advancements in each individual agent.

📄 PDF Abstract BibTeX arXiv:2404.18598

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityImage GenerationImage InpaintingLanguage ModelingLanguage ModellingLarge Language Model

Methods 이 논문이 사용한 방법론

Inpainting Train a convolutional neural network to generate the contents of an arbitrary image region conditioned on its surroundings.
Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Focus Anywhere for Fine-grained Multi-page Document Understanding

2024-05-23 · Chenglong Liu, Haoran Wei, Jinyue Chen, Lingyu Kong 외

Modern LVLMs still struggle to achieve fine-grained document understanding, such as OCR/translation/caption for regions of interest to the user, tasks that require the context of the entire page, or even multiple pages. …

document understandingOptical Character Recognition (OCR)

SOON: Scenario Oriented Object Navigation with Graph-based Exploration

2021-03-31 · CVPR 2021 1 · Fengda Zhu, Xiwen Liang, Yi Zhu, Xiaojun Chang 외

The ability to navigate like a human towards a language-guided target from anywhere in a 3D embodied environment is one of the 'holy grail' goals of intelligent robots. Most visual navigation benchmarks, however, focus o…

AttributeNavigateObjectVisual Navigation

Gaze Target Estimation Anywhere with Concepts

2026-08-11 · Xu Cao, Houze Yang, Vipin Gunda, Zhongyi Zhou 외 hf

Estimating human gaze targets from images in-the-wild is an important and formidable task. Existing approaches primarily employ brittle, multi-stage pipelines that require explicit inputs, like head bounding boxes and hu…

Gaze Target EstimationGaze Estimation

AnywhereVLA: Language-Conditioned Exploration and Mobile Manipulation

2025-09-25 · Konstantin Gubernatorov, Artem Voronov, Roman Voronov, Sergei Pasynkov 외 arxiv

We address natural language pick-and-place in unseen, unpredictable indoor environments with AnywhereVLA, a modular framework for mobile manipulation. A user text prompt serves as an entry point and is parsed into a stru…

Synthesizing Anyone, Anywhere, in Any Pose

2023-04-06 · Håkon Hukkelås, Frank Lindseth

We address the task of in-the-wild human figure synthesis, where the primary goal is to synthesize a full body given any region in any image. In-the-wild human figure synthesis has long been a challenging and under-explo…