paper-with-me

홈 › Papers

When Robots Should Say "I Don't Know": Benchmarking Abstention in Embodied Question Answering

2025-12-04 · Tao Wu, Chuhao Zhou, Guangyu Zhao, Haozhi Cao, Yewen Pu, Jianfei Yang arxiv

Embodied Question Answering (EQA) requires an agent to interpret language, perceive its environment, and navigate within 3D scenes to produce responses. Existing EQA benchmarks assume that every question must be answered, but embodied agents should know when they do not have sufficient information to answer. In this work, we focus on a minimal requirement for EQA agents, abstention: knowing when to withhold an answer. From an initial study of 500 human queries, we find that 32.4% contain missing or underspecified context. Drawing on this initial study and cognitive theories of human communication errors, we derive five representative categories requiring abstention: actionability limitation, referential underspecification, preference dependence, information unavailability, and false presupposition. We augment OpenEQA by having annotators transform well-posed questions into ambiguous variants outlined by these categories. The resulting dataset, AbstainEQA, comprises 1,636 annotated abstention cases paired with 1,636 original OpenEQA instances for balanced evaluation. Evaluating on AbstainEQA, we find that even the best frontier model only attains 42.79% abstention recall, while humans achieve 91.17%. We also find that scaling, prompting, and reasoning only yield marginal gains, and that fine-tuned models overfit to textual cues. Together, these results position abstention as a fundamental prerequisite for reliable interaction in embodied settings and as a necessary basis for effective clarification.

📄 PDF Abstract BibTeX arXiv:2512.04597

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Agentic Abstention: Do Agents Know When to Stop Instead of Act?

2026-06-27 · Han Luo, Bingbing Wen, Lucy Lu Wang arxiv

LLM agents are expected to act over multiple turns, using search, browsing interfaces, and terminal tools to complete user goals. Yet not every goal is well specified or achievable in the available environment. In such c…

Question Answering

Knowing When Not to Predict: Self Supervised Learning and Abstention for Safer DR Screening

2026-05-18 · Muskaan Chopra, Lorenz Sparrenberg, Jan H. Terheyden, Rafet Sifa arxiv

Self-supervised learning (SSL) is now a standard way to pretrain medical image models, but performance is still mostly judged by downstream accuracy. For safety-critical screening tasks such as diabetic retinopathy gradi…

Diabetic Retinopathy GradingSelf-Supervised Learning

Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded Reasoning

2026-08-28 · Vy Nguyen, Ziqi Xu, Jeffrey Chan, Estrid He 외 arxiv

Knowledge-intensive reasoning requires Large Language Models (LLMs) to ground answers in provided evidence. When evidence is insufficient, it is desirable that models abstain rather than confidently generating unsupporte…

The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents

2026-05-19 · Doguhan Yeke, Elif Su Temirel, Ananth Shreekumar, Brandon Lee 외 arxiv

Vision-language models (VLMs) are used as high-level planners for embodied agents, translating natural language instructions and visual observations into action plans. While prior work has studied abstention in LLMs, exi…

Visual Grounding

RealBirdID: Benchmarking Bird Species Identification in the Era of MLLMs

2026-03-27 · Logan Lawrence, Mustafa Chasmai, Rangel Daroya, Wuao Liu 외 arxiv

Fine-grained bird species identification in the wild is frequently unanswerable from a single image: key cues may be non-visual (e.g. vocalization), or obscured due to occlusion, camera angle, or low resolution. Yet toda…