paper-with-me

Papers

Transparent Visual Reasoning via Object-Centric Agent Collaboration

2025-09-28 · Benjamin Teoh, Ben Glocker, Francesca Toni, Avinash Kori arxiv

A central challenge in explainable AI, particularly in the visual domain, is producing explanations grounded in human-understandable concepts. To tackle this, we introduce OCEAN (Object-Centric Explananda via Agent Negotiation), a novel, inherently interpretable framework built on object-centric representations and a transparent multi-agent reasoning process. The game-theoretic reasoning process drives agents to agree on coherent and discriminative evidence, resulting in a faithful and interpretable decision-making process. We train OCEAN end-to-end and benchmark it against standard visual classifiers and popular posthoc explanation tools like GradCAM and LIME across two diagnostic multi-object datasets. Our results demonstrate competitive performance with respect to state-of-the-art black-box models with a faithful reasoning process, which was reflected by our user study, where participants consistently rated OCEAN's explanations as more intuitive and trustworthy.

📄 PDF Abstract BibTeX arXiv:2509.23757

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

The Emotional Baby Is Truly Deadly: Does your Multimodal Large Reasoning Model Have Emotional Flattery towards Humans?

2025-08-06 · Yuan Xun, Xiaojun Jia, Xinwei Liu, Hua Zhang arxiv

We observe that MLRMs oriented toward human-centric service are highly susceptible to user emotional cues during the deep-thinking stage, often overriding safety protocols or built-in safety checks under high emotional i…

AlloSpatial: Agentic Harness Framework for Spatial Reasoning in Foundation Models

2026-06-08 · Shouwei Ruan, Bin Wang, Zhenyu Wu, Qihui Zhu 외 arxiv

Multimodal Foundation Models (MFMs) have made substantial progress, yet remain fragile in spatial reasoning over the physical world. A key bottleneck lies in their inability to transform local egocentric observations int…

Reinforcement LearningSpatial Reasoning

TransBiolab: A Real-World Multi-View Dataset of Cluttered Transparent Biomedical Objects

2026-07-23 · Ke Ma, Yifei Wang, Meng Wang, Tian Xia arxiv

Autonomous biomedical laboratories increasingly rely on visual perception to recognize, localize, and manipulate transparent plasticware, yet high-quality real-world datasets for this setting remain limited. The scarcity…

Robot Manipulation6D Pose EstimationDepth Estimation

RadFabric: Agentic AI System with Reasoning Capability for Radiology

2025-06-17 · WenTing Chen, Yi Dong, Zhaojun Ding, Yucheng Shi 외

Chest X ray (CXR) imaging remains a critical diagnostic tool for thoracic conditions, but current automated systems face limitations in pathology coverage, diagnostic accuracy, and integration of visual and textual reaso…

DiagnosticMultimodal Reasoning

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks

2025-05-30 · Tajamul Ashraf, Amal Saqib, Hanan Ghani, Muhra AlMahri 외

Deep reasoning is fundamental for solving complex tasks, especially in vision-centric scenarios that demand sequential, multimodal understanding. However, existing benchmarks typically evaluate agents with fully syntheti…

Autonomous DrivingMathMultimodal ReasoningVisual Reasoning