paper-with-me

홈 › Papers

Unsolvable Problem Detection: Evaluating Trustworthiness of Vision Language Models

2024-03-29 · Atsuyuki Miyai, Jingkang Yang, Jingyang Zhang, Yifei Ming, Qing Yu, Go Irie, Yixuan Li, Hai Li, Ziwei Liu, Kiyoharu Aizawa

This paper introduces a novel and significant challenge for Vision Language Models (VLMs), termed Unsolvable Problem Detection (UPD). UPD examines the VLM's ability to withhold answers when faced with unsolvable problems in the context of Visual Question Answering (VQA) tasks. UPD encompasses three distinct settings: Absent Answer Detection (AAD), Incompatible Answer Set Detection (IASD), and Incompatible Visual Question Detection (IVQD). To deeply investigate the UPD problem, extensive experiments indicate that most VLMs, including GPT-4V and LLaVA-Next-34B, struggle with our benchmarks to varying extents, highlighting significant room for the improvements. To address UPD, we explore both training-free and training-based solutions, offering new insights into their effectiveness and limitations. We hope our insights, together with future efforts within the proposed UPD settings, will enhance the broader understanding and development of more practical and reliable VLMs.

📄 PDF Abstract BibTeX arXiv:2403.20331

Code (1)

atsumiyai/upd 공식 구현 pytorch

Tasks

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Learning the Boundary of Solvability: Aligning LLMs to Detect Unsolvable Problems

2025-12-01 · Dengyun Peng, Qiguang Chen, Bofei Liu, Jiannan Guan 외 arxiv

Ensuring large language model (LLM) reliability requires distinguishing objective unsolvability (inherent contradictions) from subjective capability limitations (tasks exceeding model competence). Current LLMs often conf…

Reinforcement Learning

Multi-Task Ordinal Regression for Jointly Predicting the Trustworthiness and the Leading Political Ideology of News Media

2019-04-01 · NAACL 2019 6 · Ramy Baly, Georgi Karadzhov, Abdelrhman Saleh, James Glass 외

In the context of fake news, bias, and propaganda, we study two important but relatively under-explored problems: (i) trustworthiness estimation (on a 3-point scale) and (ii) political ideology detection (left/right bias…

Articles

ReliableMath: Benchmark of Reliable Mathematical Reasoning on Large Language Models

2025-07-03 · Boyang Xue, Qi Zhu, Rui Wang, Sheng Wang 외 arxiv

Although demonstrating remarkable performance on reasoning tasks, Large Language Models (LLMs) still tend to fabricate unreliable responses when confronted with problems that are unsolvable or beyond their capability, se…

Mathematical Reasoning

Planning Task Shielding: Detecting and Repairing Flaws in Planning Tasks through Turning them Unsolvable

2026-04-08 · Alberto Pozanco, Marianela Morales, Pietro Totis, Daniel Borrajo arxiv

Most research in planning focuses on generating a plan to achieve a desired set of goals. However, a goal specification can also be used to encode a property that should never hold, allowing a planner to identify a trace…

Evaluating Trustworthiness of Online News Publishers via Article Classification

2024-01-03 · John Bianchi, Manuel Pratelli, Marinella Petrocchi, Fabio Pinelli

The proliferation of low-quality online information in today's era has underscored the need for robust and automatic mechanisms to evaluate the trustworthiness of online news publishers. In this paper, we analyse the tru…

ArticlesClassification