paper-with-me

Papers

Bridging Human and LLM Judgments: Understanding and Narrowing the Gap

2025-08-18 · Felipe Maia Polo, Xinhe Wang, Mikhail Yurochkin, Gongjun Xu, Moulinath Banerjee, Yuekai Sun arxiv

Large language models are increasingly used as judges (LLM-as-a-judge) to evaluate model outputs at scale, but their assessments often diverge systematically from human judgments. We present Bridge, a unified statistical framework that explicitly bridges human and LLM evaluations under both absolute scoring and pairwise comparison paradigms. Bridge posits a latent human preference score for each prompt-response pair and models LLM deviations as linear transformations of covariates that capture sources of discrepancies. This offers a simple and principled framework for refining LLM ratings and characterizing systematic discrepancies between humans and LLMs. We provide an efficient fitting algorithm with asymptotic guarantees for statistical inference. Using six LLM judges and two benchmarks (BigGen Bench and Chatbot Arena), Bridge achieves higher agreement with human ratings (accuracy, calibration, and KL divergence) and exposes systematic human-LLM gaps.

📄 PDF Abstract BibTeX arXiv:2508.12792

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Human Semantic Representations of Social Interactions from Moving Shapes

2025-09-25 · Yiling Yun, Hongjing Lu arxiv

Humans are social creatures who readily recognize various social interactions from simple display of moving shapes. While previous research has often focused on visual features, we examine what semantic representations t…

Enhancing Target-Guided Proactive Dialogue Systems via Conversational Scenario Modeling and Intent-Keyword Bridging

2026-05-12 · Maodong Li, Yancui Li, Fang Kong arxiv

A target-guided proactive dialogue system aims to steer conversations proactively toward pre-defined targets, such as designated keywords or specific topics. During guided conversations, dynamically modeling conversation…

Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments

2024-06-17 · Han Zhou, Xingchen Wan, Yinhong Liu, Nigel Collier 외

Large language models (LLMs) have shown promising abilities as cost-effective and reference-free evaluators for assessing language generation quality. In particular, pairwise LLM evaluators, which compare two generated t…

FairnessLanguage ModelingLanguage ModellingLarge Language Model+2

SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation

2025-11-21 · Shrikant Kendre, Austin Xu, Honglu Zhou, Michael Ryoo 외 arxiv

Traditional evaluation metrics for textual and visual question answering, like ROUGE, METEOR, and Exact Match (EM), focus heavily on n-gram based lexical similarity, often missing the deeper semantic understanding needed…

Visual Question Answering

Learning to see people like people

2017-05-05 · Amanda Song, Linjie Li, Chad Atalla, Garrison Cottrell

Humans make complex inferences on faces, ranging from objective properties (gender, ethnicity, expression, age, identity, etc) to subjective judgments (facial attractiveness, trustworthiness, sociability, friendliness, e…