paper-with-me

Papers

Extending Minimal Pairs with Ordinal Surprisal Curves and Entropy Across Applied Domains

2026-03-15 · Andrew Katz arxiv

The minimal pairs paradigm of comparing model probabilities for contrasting completions has proven useful for evaluating linguistic knowledge in language models, yet its application has largely been confined to binary grammaticality judgments over syntactic phenomena. Additionally, standard prompting-based evaluation requires expensive text generation, may elicit post-hoc rationalizations rather than model judgments, and discards information about model uncertainty. We address both limitations by extending surprisal-based evaluation from binary grammaticality contrasts to ordinal-scaled classification and scoring tasks across multiple domains. Rather than asking models to generate answers, we measure the information-theoretic "surprise" (negative log probability) they assign to each position on rating scales (e.g., 1-5 or 1-9), yielding full surprisal curves that reveal both the model's preferred response and its uncertainty via entropy. We explore this framework across four domains: social-ecological-technological systems classification, causal statement identification (binary and scaled), figurative language detection, and deductive qualitative coding. Across these domains, surprisal curves produce interpretable classification signals with clear minima near expected ordinal scale positions, and entropy over the completion tended to distinguish genuinely ambiguous items from easier items.

📄 PDF Abstract BibTeX arXiv:2603.14400

Code (0)

등록된 구현이 없습니다.

Tasks

Text Generation

Similar Papers 제목 키워드 기반

Surprisal-Guided Selection: Compute-Optimal Test-Time Strategies for Execution-Grounded Code Generation

2026-02-07 · Jarrod Barnes arxiv

Test-time training (TTT) adapts language models through gradient-based updates at inference. But is adaptation the right strategy? We study compute-optimal test-time strategies for verifiable execution-grounded (VEG) tas…

Code Generation

Characterizing Learning Curves During Language Model Pre-Training: Learning, Forgetting, and Stability

2023-08-29 · Tyler A. Chang, Zhuowen Tu, Benjamin K. Bergen

How do language models learn to make predictions during pre-training? To study this, we extract learning curves from five autoregressive English language model pre-training runs, for 1M unseen tokens in context. We obser…

Language ModelingLanguage Modelling

Controlling Surprisal in Music Generation via Information Content Curve Matching

2024-08-12 · Mathias Rose Bjare, Stefan Lattner, Gerhard Widmer

In recent years, the quality and public interest in music generation systems have grown, encouraging research into various ways to control these systems. We propose a novel method for controlling surprisal in music gener…

Music Generation

VORD: Visual Ordinal Calibration for Mitigating Object Hallucinations in Large Vision-Language Models

2024-12-20 · Dexter Neo, Tsuhan Chen

Large Vision-Language Models (LVLMs) have made remarkable developments along with the recent surge of large language models. Despite their advancements, LVLMs have a tendency to generate plausible yet inaccurate or incon…

Investigating the Utility of Surprisal from Large Language Models for Speech Synthesis Prosody

2023-06-16 · Sofoklis Kakouros, Juraj Šimko, Martti Vainio, Antti Suni

This paper investigates the use of word surprisal, a measure of the predictability of a word in a given context, as a feature to aid speech synthesis prosody. We explore how word surprisal extracted from large language m…

Speech Synthesis