paper-with-me

홈 › Papers

The Partial Testimony of Logs: Evaluation of Language Model Generation under Confounded Model Choice

2026-05-02 · Jikai Jin, Vasilis Syrgkanis arxiv

Offline evaluation of language models from usage logs is biased when model choice is confounded: the same user-side factors that influence which model is used can also influence how its output is judged, so raw comparisons of logged scores mix self-selected populations rather than estimating a common quantity of interest. A small randomized experiment can break this bias by overriding model choice, but in practice such experiments are scarce and costly. We study a three-source design that combines a large confounded observational log (OBS) for scale, a small randomized experiment (EXP) for unconfounded scoring, and an offline simulator (SIM) that replays candidate models on cached contexts. Our main result is an identification theorem showing that the randomized experiment and the simulator are together enough to recover causal model values; the observational log enters only afterward, to reduce estimation error rather than to make the causal comparison valid. Six estimator families are evaluated in a controlled semi-synthetic validation and in two real-task cached benchmarks for summarization and coding. No family dominates every regime; relative performance depends on the amount of unbiased EXP supervision and on how closely the target reward aligns with OBS-derived structure.

📄 PDF Abstract BibTeX arXiv:2605.01311

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Specifying and Staging Mixed-Initiative Dialogs with Program Generation and Transformation

2011-08-02 · Saverio Perugini

Specifying and implementing flexible human-computer dialogs, such as those used in kiosks and smart phone apps, is challenging because of the numerous and varied directions in which each user might steer a dialog. The ob…

Interpretive Blindness

2021-10-19 · Nicholas Asher, Julie Hunter

We model here an epistemic bias we call \textit{interpretive blindness} (IB). IB is a special problem for learning from testimony, in which one acquires information only from text or conversation. We show that IB follows…

The Inconsistency Critique: Epistemic Practices and AI Testimony About Inner States

2025-12-22 · Gerol Petruzella arxiv

The question of whether AI systems have morally relevant interests -- the 'model welfare' question -- depends in part on how we evaluate AI testimony about inner states. This paper develops what I call the inconsistency …

Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation

2026-08-13 · Divya Vetticaden, Arya Gupta, Julian Nyarko, Megan Ma arxiv

Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simu…

New Dimensions in Testimony Demonstration

2016-06-01 · NAACL 2016 6 · Ron Artstein, Alesia Gainer, Kallirroi Georgila, Anton Leuski 외