paper-with-me

홈 › Papers

Marmot: Multi-Agent Reasoning for Multi-Object Self-Correcting in Improving Image-Text Alignment

2025-04-10 · Jiayang Sun, Hongbo Wang, Jie Cao, Huaibo Huang, Ran He

While diffusion models excel at generating high-quality images, they often struggle with accurate counting, attributes, and spatial relationships in complex multi-object scenes. One potential approach is to utilize Multimodal Large Language Model (MLLM) as an AI agent to build a self-correction framework. However, these approaches are highly dependent on the capabilities of the employed MLLM, often failing to account for all objects within the image. To address these challenges, we propose Marmot, a novel and generalizable framework that employs Multi-Agent Reasoning for Multi-Object Self-Correcting, enhancing image-text alignment and facilitating more coherent multi-object image editing. Our framework adopts a divide-and-conquer strategy, decomposing the self-correction task into object-level subtasks according to three critical dimensions: counting, attributes, and spatial relationships. We construct a multi-agent self-correcting system featuring a decision-execution-verification mechanism, effectively mitigating inter-object interference and enhancing editing reliability. To resolve the problem of subtask integration, we propose a Pixel-Domain Stitching Smoother that employs mask-guided two-stage latent space optimization. This innovation enables parallel processing of subtask results, thereby enhancing runtime efficiency while eliminating multi-stage distortion accumulation. Extensive experiments demonstrate that Marmot significantly improves accuracy in object counting, attribute assignment, and spatial relationships for image generation tasks.

📄 PDF Abstract BibTeX arXiv:2504.20054

Code (0)

등록된 구현이 없습니다.

Tasks

AI AgentAttributeImage GenerationLarge Language ModelMultimodal Large Language ModelObjectObject Counting

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

MARMOT: Masked Autoencoder for Modeling Transient Imaging

2025-06-10 · Siyuan Shen, Ziheng Wang, Xingyue Peng, Suan Xia 외

Pretrained models have demonstrated impressive success in many modalities such as language and vision. Recent works facilitate the pretraining paradigm in imaging research. Transients are a novel modality, which are capt…

Decoder

MARMOT: A Deep Learning Framework for Constructing Multimodal Representations for Vision-and-Language Tasks

2021-09-23 · Patrick Y. Wu, Walter R. Mebane Jr

Political activity on social media presents a data-rich window into political behavior, but the vast amount of data means that almost all content analyses of social media require a data labeling step. However, most autom…

Translation

MARMOT: A Toolkit for Translation Quality Estimation at the Word Level

2016-05-01 · LREC 2016 5 · Varvara Logacheva, Chris Hokamp, Lucia Specia

We present Marmot{\textasciitilde}― a new toolkit for quality estimation (QE) of machine translation output. Marmot contains utilities targeted at quality estimation at the word and phrase level. However, due to its fl…

Machine TranslationSentenceTranslation

COIN: Collaborative Interaction-Aware Multi-Agent Reinforcement Learning for Self-Driving Systems

2026-03-26 · Yifeng Zhang, Jieming Chen, Tingguang Zhou, Tanishq Duhan 외 arxiv

Multi-Agent Self-Driving (MASD) systems provide an effective solution for coordinating autonomous vehicles to reduce congestion and enhance both safety and operational efficiency in future intelligent transportation syst…

Multi-agent Reinforcement LearningAutonomous Vehicles

Synthetic Data Augmentation for Table Detection: Re-evaluating TableNet's Performance with Automatically Generated Document Images

2025-06-17 · Krishna Sahukara, Zineddine Bettouche, Andreas Fischer

Document pages captured by smartphones or scanners often contain tables, yet manual extraction is slow and error-prone. We introduce an automated LaTeX-based pipeline that synthesizes realistic two-column pages with visu…

Data AugmentationTable Detection