paper-with-me

Papers

Explicit Abstention Knobs for Predictable Reliability in Video Question Answering

2025-12-31 · Jorge Ortiz arxiv

High-stakes deployment of vision-language models (VLMs) requires selective prediction, where systems abstain when uncertain rather than risk costly errors. We investigate whether confidence-based abstention provides reliable control over error rates in video question answering, and whether that control remains robust under distribution shift. Using NExT-QA and Gemini 2.0 Flash, we establish two findings. First, confidence thresholding provides mechanistic control in-distribution. Sweeping threshold epsilon produces smooth risk-coverage tradeoffs, reducing error rates f

📄 PDF Abstract BibTeX arXiv:2601.00138

Code (0)

등록된 구현이 없습니다.

Tasks

Video Question Answering

Similar Papers 제목 키워드 기반

RAS: a Reliability Oriented Metric for Automatic Speech Recognition

2026-04-27 · Wenbin Huang, Yuhang Qiu, Bohan Li, Yiwei Guo 외 arxiv

Automatic speech recognition systems often produce confident yet incorrect transcriptions under noisy or ambiguous conditions, which can be misleading for both users and downstream applications. Standard evaluation based…

Reinforcement LearningSpeech Recognition

Sacred or Synthetic? Evaluating LLM Reliability and Abstention for Religious Questions

2025-08-04 · Farah Atif, Nursultan Askarbekuly, Kareem Darwish, Monojit Choudhury arxiv

Despite the increasing usage of Large Language Models (LLMs) in answering questions in a variety of domains, their reliability and accuracy remain unexamined for a plethora of domains including the religious domains. In …

Boundary-Aware NL2SQL: Integrating Reliability through Hybrid Reward and Data Synthesis

2026-01-15 · Songsong Tian, Kongsheng Zhuo, Zhendong Wang, Rong Shen 외 arxiv

In this paper, we present BAR-SQL (Boundary-Aware Reliable NL2SQL), a unified training framework that embeds reliability and boundary awareness directly into the generation process. We introduce a Seed Mutation data synt…

Reinforcement Learning

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions

2025-06-10 · Polina Kirichenko, Mark Ibrahim, Kamalika Chaudhuri, Samuel J. Bell

For Large Language Models (LLMs) to be reliably deployed in both everyday and high-stakes domains, knowing when not to answer is equally critical as answering correctly. Real-world user queries, which can be underspecifi…

Math

Rewarding Intellectual Humility Learning When Not To Answer In Large Language Models

2026-01-27 · Abha Jha, Akanksha Mahajan, Ashwath Vaithinathan Aravindan, Praveen Saravanan 외 arxiv

Large Language Models (LLMs) often produce hallucinated or unverifiable content, undermining their reliability in factual domains. This work investigates Reinforcement Learning with Verifiable Rewards (RLVR) as a trainin…

Reinforcement LearningQuestion Answering