paper-with-me

홈 › Papers

What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models

2026-01-07 · Dasol Choi, Guijin Son, Hanwool Lee, Minhyuk Kim, Hyunwoo Ko, Teabin Lim, Ahn Eungyeol, Jungwhan Kim, Seunghyeok Hong, Youngsook Song arxiv

Current vision-language benchmarks predominantly feature well-structured questions with clear, explicit prompts. However, real user queries are often informal and underspecified. Users naturally leave much unsaid, relying on images to convey context. We introduce HAERAE-Vision, a benchmark of 653 real-world visual questions from Korean online communities (0.76% survival from 86K candidates), each paired with an explicit rewrite, yielding 1,306 query variants in total. Evaluating 39 VLMs, we find that even state-of-the-art models (GPT-5, Gemini 2.5 Pro) achieve under 50% on the original queries. Crucially, query explicitation alone yields 8 to 22 point improvements, with smaller models benefiting most. We further show that even with web search, under-specified queries underperform explicit queries without search, revealing that current retrieval cannot compensate for what users leave unsaid. Our findings demonstrate that a substantial portion of VLM difficulty stem from natural query under-specification instead of model capability, highlighting a critical gap between benchmark evaluation and real-world deployment.

📄 PDF Abstract BibTeX arXiv:2601.06165

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RealFin: How Well Do LLMs Reason About Finance When Users Leave Things Unsaid?

2026-02-06 · Yuyang Dai, Yan Lin, Zhuohan Xie, Yuxia Wang arxiv

Reliable financial reasoning requires knowing not only how to answer, but also when an answer cannot be justified. In real financial practice, problems often rely on implicit assumptions that are taken for granted rather…

Do Pre-Trained Language Models Detect and Understand Semantic Underspecification? Ask the DUST!

2024-02-19 · Frank Wildenburg, Michael Hanna, Sandro Pezzelle

In everyday language use, speakers frequently utter and interpret sentences that are semantically underspecified, namely, whose content is insufficient to fully convey their message or interpret them univocally. For exam…

Sentence

STaR-GATE: Teaching Language Models to Ask Clarifying Questions

2024-03-28 · Chinmaya Andukuri, Jan-Philipp Fränken, Tobias Gerstenberg, Noah D. Goodman

When prompting language models to complete a task, users often leave important aspects unsaid. While asking questions could resolve this ambiguity (GATE; Li et al., 2023), models often struggle to ask good questions. We …

Language ModelingLanguage Modelling

Text as Partial Constraint: Core-Residual Alignment for Robust Vision-Language Learning

2026-07-03 · Chengzhen Yu, Canran Xiao, Siyuan Ma, Yang Liu arxiv

Vision-language alignment powers open-vocabulary recognition, retrieval, and LVLM grounding, yet natural captions are often underspecified, making similarity brittle and overly confident under paraphrase and omitted deta…

Adversarial Robustness

Beyond expert users: agents should help users construct preferences, not just elicit them

2026-06-29 · Irena Saracay, Ludwig Schmidt, Carlos Guestrin arxiv

Agents typically assume an expert user -- one with well-formed preferences about what they want -- and default to clarifying questions whenever the task is underspecified. We argue this assumption is unrealistic. Users o…