paper-with-me

Papers

Blind to the Human Touch: Overlap Bias in LLM-Based Summary Evaluation

2026-02-07 · Jiangnan Fang, Cheng-Tse Liu, Hanieh Deilamsalehy, Nesreen K. Ahmed, Puneet Mathur, Nedim Lipka, Franck Dernoncourt, Ryan A. Rossi arxiv

Large language model (LLM) judges have often been used alongside traditional, algorithm-based metrics for tasks like summarization because they better capture semantic information, are better at reasoning, and are more robust to paraphrasing. However, LLM judges show biases for length and order among others, and are vulnerable to various adversarial input prompts. While recent studies have looked into these biases, few have analyzed them at a more granular level in relation to a well-defined overlap metric. In this work we provide an LLM judge bias analysis as a function of overlap with human-written responses in the domain of summarization. We test 9 recent LLMs with parameter counts ranging from 1 billion to 12 billion, including variants of Gemma 3 and LLaMA 3. We find that LLM judges increasingly prefer summaries generated by other LLMs over those written by humans as the similarities (as measured by ROUGE and BLEU) between the judged summaries decrease, and this pattern extends to all but one model tested, and exists regardless of the models' own position biases. Additionally, we find that models struggle to judge even summaries with limited overlaps, suggesting that LLM-as-a-judge in the summary domain should rely on techniques beyond a simple comparison.

📄 PDF Abstract BibTeX arXiv:2602.07673

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Text to Blind Motion

2024-12-06 · Hee Jae Kim, Kathakoli Sengupta, Masaki Kuribayashi, Hernisa Kacorri 외

People who are blind perceive the world differently than those who are sighted, which can result in distinct motion characteristics. For instance, when crossing at an intersection, blind individuals may have different pa…

Autonomous Vehicles

Spot the BlindSpots: Systematic Identification and Quantification of Fine-Grained LLM Biases in Contact Center Summaries

2025-08-18 · Kawin Mayilvaghanan, Siddhant Gupta, Ayush Kumar arxiv

Abstractive summarization is a core application in contact centers, where Large Language Models (LLMs) generate millions of summaries of call transcripts daily. Despite their apparent quality, it remains unclear whether …

Summaformers @ LaySumm 20, LongSumm 20

2021-01-10 · Sayar Ghosh Roy, Nikhil Pinnaparaju, Risubh Jain, Manish Gupta 외

Automatic text summarization has been widely studied as an important task in natural language processing. Traditionally, various feature engineering and machine learning based systems have been proposed for extractive as…

Abstractive Text SummarizationFeature EngineeringText Summarization

A Geometric Approach For Fully Automatic Chromosome Segmentation

2011-12-18 · Shervin Minaee, Mehran Fotouhi, Babak Hossein Khalaj

A fundamental task in human chromosome analysis is chromosome segmentation. Segmentation plays an important role in chromosome karyotyping. The first step in segmentation is to remove intrusive objects such as stain debr…

Segmentation

StateLens: A Reverse Engineering Solution for Making Existing Dynamic Touchscreens Accessible

2019-08-20 · Anhong Guo, Junhan Kong, Michael Rivera, Frank F. Xu 외

Blind people frequently encounter inaccessible dynamic touchscreens in their everyday lives that are difficult, frustrating, and often impossible to use independently. Touchscreens are often the only way to control every…