paper-with-me

홈 › Papers

Calibrating the Subjective

2024-12-24 · Mark Whitmeyer

I conduct Rabin's (2000) calibration exercise in the subjective expected utility realm. I show that the rejection of some risky bet by a risk-averse agent only implies the rejection of more extreme and less desirable bets and nothing more.

📄 PDF Abstract BibTeX arXiv:2412.18486

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Crowd-Calibrator: Can Annotator Disagreement Inform Calibration in Subjective Tasks?

2024-08-26 · Urja Khurana, Eric Nalisnick, Antske Fokkens, Swabha Swayamdipta

Subjective tasks in NLP have been mostly relegated to objective standards, where the gold label is decided by taking the majority vote. This obfuscates annotator disagreement and the inherent uncertainty of the label. We…

Decision MakingHate Speech DetectionNatural Language Inference

Judging with Confidence: Calibrating Autoraters to Preference Distributions

2025-09-30 · Zhuohang Li, Xiaowei Li, Chengyu Huang, Guowang Li 외 arxiv

The alignment of large language models (LLMs) with human values increasingly relies on using other LLMs as automated judges, or ``autoraters''. However, their reliability is limited by a foundational issue: they are trai…

Reinforcement Learning

Calibrating “Cheap Signals” in Peer Review without a Prior

2023-09-21 · NeurIPS 2023 11

Peer review lies at the core of the academic process, but even well-intentioned reviewers can still provide noisy ratings. While ranking papers by average ratings may reduce noise, varying noise levels and systematic bia…

Semantic Latent Space Regression of Diffusion Autoencoders for Vertebral Fracture Grading

2023-03-21 · Matthias Keicher, Matan Atad, David Schinz, Alexandra S. Gersing 외

Vertebral fractures are a consequence of osteoporosis, with significant health implications for affected patients. Unfortunately, grading their severity using CT exams is hard and subjective, motivating automated grading…

regression

Calibrating LLM Judges: Linear Probes for Fast and Reliable Uncertainty Estimation

2025-12-23 · Bhaktipriya Radharapu, Eshika Saxena, Kenneth Li, Chenxi Whitehouse 외 arxiv

As LLM-based judges become integral to industry applications, obtaining well-calibrated uncertainty estimates efficiently has become critical for production deployment. However, existing techniques, such as verbalized co…