paper-with-me

홈 › Papers

Generation Space Size: Understanding and Calibrating Open-Endedness of LLM Generations

2025-10-14 · Sunny Yu, Ahmad Jabbar, Robert Hawkins, Dan Jurafsky, Myra Cheng arxiv

Different open-ended generation tasks require different degrees of output diversity. However, current LLMs are often miscalibrated. They collapse to overly homogeneous outputs for creative tasks and hallucinate diverse but incorrect responses for factual tasks. We argue that these two failure modes are unified by, and can both be addressed by, the notion of effective generation space size (GSS) -- the set of semantically distinct outputs a model considers for a prompt. We present GSSBench, a task suite of prompt pairs with ground-truth GSS relationships to assess different metrics and understand where models diverge from desired behavior. We find that hallucination detection metrics, particularly EigenScore, consistently outperform standard diversity and uncertainty quantification metrics, while using only model internals, providing interpretable insights into a model's internal task representations. We demonstrate three applications of GSS: (1) detecting prompt ambiguity and predicting clarification questions for better grounding, (2) interpreting overthinking and underthinking in reasoning models, and (3) steering models to expand their generation space to yield high-quality and diverse outputs.

📄 PDF Abstract BibTeX arXiv:2510.12699

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Know Your Space: Inlier and Outlier Construction for Calibrating Medical OOD Detectors

2022-07-12 · Vivek Narayanaswamy, Yamen Mubarka, Rushil Anirudh, Deepta Rajan 외

We focus on the problem of producing well-calibrated out-of-distribution (OOD) detectors, in order to enable safe deployment of medical image classifiers. Motivated by the difficulty of curating suitable calibration data…

Data AugmentationOpen Set LearningOut-of-Distribution DetectionOut of Distribution (OOD) Detection+1

Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs

2026-07-07 · Andrea Alfarano, Andrea Bacciu, Saab Mansour, Amin Mantrach 외 arxiv

Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English. We present the first large-scale evaluation of UE methods across 22 langua…

Interactive introduction to self-calibrating interfaces

2022-12-12 · Jonathan Grizou

This interactive paper aims to provide an intuitive understanding of the self-calibrating interface paradigm. Under this paradigm, you can choose how to use an interface which can adapt to your preferences on the fly. We…

Improving Fabrication Fidelity of Integrated Nanophotonic Devices Using Deep Learning

2023-03-21 · Dusan Gostimirovic, Yuri Grinberg, Dan-Xia Xu, Odile Liboiron-Ladouceur

Next-generation integrated nanophotonic device designs leverage advanced optimization techniques such as inverse design and topology optimization which achieve high performance and extreme miniaturization by optimizing a…

Deep Learning

Calibrating Sequence likelihood Improves Conditional Language Generation

2022-09-30 · Yao Zhao, Misha Khalman, Rishabh Joshi, Shashi Narayan 외

Conditional language models are predominantly trained with maximum likelihood estimation (MLE), giving probability mass to sparsely observed target sequences. While MLE trained models assign high probability to plausible…

abstractive question answeringAbstractive Text SummarizationBlockingData-to-Text Generation+5