paper-with-me

Papers

Seeing Before Agreeing: Aligning Multi-Agent Consensus with Visual Evidence

2026-05-29 · Yuhan Wang, Shuochen Chang, Yalin Feng, Dongsheng Ma, Yuanzi Li, Zhengren Wang, Yinglong Yang, Yufei Chen, Yikang Wang, Shaoxu Sun, Wentao Zhang arxiv

Vision-language models (VLMs) have achieved strong performance on visual question answering (VQA). To mitigate individual hallucinations and blind spots, aggregating diverse perspectives via multi-agent collaboration has emerged as a promising paradigm. While this approach has shown great success in textual QA, its potential in the multimodal domain remains under-explored. Existing multi-agent VQA methods predominantly adapt text-centric protocols, focusing on textual discussions while ignoring the alignment of visual information. In this work, we reveal a key insight: answer-level agreement is insufficient for reliable multi-agent VQA; \textit{aligned visual evidence} -- shared support from the image regions agents rely on -- is essential for trustworthy consensus. To leverage this insight, we propose EAGLE (\textbf{E}vidence-\textbf{A}ligned \textbf{G}rounded mu\textbf{L}ti-agent r\textbf{E}asoning), a training-free evidence-centered framework for coordinating multiple VLM agents. EAGLE explicitly exposes each agent's grounding regions as visual evidence, enables mutual verification over the evidence, and uses evidence consistency to guide final decision-making. Experiments on six VQA benchmarks show that EAGLE achieves best average performance across domains while remaining lightweight, interpretable, and practical for deployment.

📄 PDF Abstract BibTeX arXiv:2605.30698

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Question Answering

Similar Papers 제목 키워드 기반

JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety

2026-07-22 · Yuan Xiong, Linji Hao, Shizhu He, Yequan Wang 외 arxiv

Agent safety is moving from content moderation toward preventing operational failures before tool-using agents act. We propose Janus, a foresight-oriented framework for long-horizon agent safety that trains guards to ant…

Seeing the Whole in the Parts in Self-Supervised Representation Learning

2025-01-06 · Arthur Aubret, Céline Teulière, Jochen Triesch

Recent successes in self-supervised learning (SSL) model spatial co-occurrences of visual features either by masking portions of an image or by aggressively cropping it. Here, we propose a new way to model spatial co-occ…

Representation LearningSelf-Supervised Learning

Sufficient Conditions on Bipartite Consensus of Weakly Connected Matrix-weighted Networks

2023-07-03 · Chongzhi Wang, Haibin Shao, Ying Tan, Dewei Li

Recent advancements in bipartite consensus, a scenario where agents are divided into two disjoint sets with agents in the same set agreeing on a certain value and those in different sets agreeing on opposite or specifica…

Responsibility in Extensive Form Games

2023-12-12 · Qi Shi

Two different forms of responsibility, counterfactual and seeing-to-it, have been extensively discussed in the philosophy and AI in the context of a single agent or multiple agents acting simultaneously. Although the gen…

counterfactualFormPhilosophy

Collaborative AI Teaming in Unknown Environments via Active Goal Deduction

2024-03-22 · Zuyuan Zhang, Hanhan Zhou, Mahdi Imani, Taeyoung Lee 외

With the advancements of artificial intelligence (AI), we're seeing more scenarios that require AI to work closely with other agents, whose goals and strategies might not be known beforehand. However, existing approaches…

StarcraftStarcraft II