paper-with-me

홈 › Papers

Who is the richest club in the championship? Detecting and Rewriting Underspecified Questions Improve QA Performance

2026-02-12 · Yunchong Huang, Gianni Barlacchi, Sandro Pezzelle arxiv

Large language models (LLMs) perform well on well-posed questions, yet standard question-answering (QA) benchmarks remain far from solved. We argue that this gap is partly due to underspecified questions - queries whose interpretation cannot be uniquely determined without additional context. To test this hypothesis, we introduce an LLM-based classifier to identify underspecified questions and apply it to several widely used QA datasets, finding that 16% to over 50% of benchmark questions are underspecified and that LLMs perform significantly worse on them. To isolate the effect of underspecification, we conduct a controlled rewriting experiment that serves as an upper-bound analysis, rewriting underspecified questions into fully specified variants while holding gold answers fixed. QA performance consistently improves under this setting, indicating that many apparent QA failures stem from question underspecification rather than model limitations. Our findings highlight underspecification as an important confound in QA evaluation and motivate greater attention to question clarity in benchmark design.

📄 PDF Abstract BibTeX arXiv:2602.11938

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

BetaRun Soccer Simulation League Team: Variety, Complexity, and Learning

2017-03-12 · Olivia Michael, Oliver Obst

RoboCup offers a set of benchmark problems for Artificial Intelligence in form of official world championships since 1997. The most tactical advanced and richest in terms of behavioural complexity of these is the 2D Socc…

BIG-bench Machine Learning

Improving Neural Retrieval with Attribution-Guided Query Rewriting

2026-02-12 · Moncef Garouani, Josiane Mothe arxiv

Neural retrievers are effective but brittle: underspecified or ambiguous queries can misdirect ranking even when relevant documents exist. Existing approaches address this brittleness only partially: LLMs rewrite queries…

RECAP: REwriting Conversations for Intent Understanding in Agentic Planning

2025-08-29 · Kushan Mitra, Dan Zhang, Hannah Kim, Estevam Hruschka arxiv

Understanding user intent is essential for effective planning in conversational assistants, particularly those powered by large language models (LLMs) coordinating multiple agents. However, real-world dialogues are often…

Intent Detection

Context Aware Query Rewriting for Text Rankers using LLM

2023-08-31 · Abhijit Anand, Venktesh V, Vinay Setty, Avishek Anand

Query rewriting refers to an established family of approaches that are applied to underspecified and ambiguous queries to overcome the vocabulary mismatch problem in document ranking. Queries are typically rewritten duri…

Document RankingPassage Ranking

Beyond Supervised Clarification: Input Rewriting with LLMs for Dialogue Discourse Parsing

2026-07-02 · Yiming Liu, Ziyue Zhang, Zhichao Xu, Xin Yu 외 arxiv

Rewriting inputs to improve frozen downstream models has become a common strategy in modern NLP pipelines. Prior work on incremental dialogue discourse parsing (DDP) shows that supervised clarification models can rewrite…

Discourse Parsing