paper-with-me

홈 › Papers

Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs

2026-02-04 · Letian Cheng, Junyan Wang, Yan Gao, Elliott Wen, Ting Dang, Hong Jia arxiv

Perplexity is a widely adopted metric for assessing the predictive quality of large language models (LLMs) and often serves as a reference metric for downstream evaluations. However, recent evidence shows that perplexity can be unreliable, especially when irrelevant long inputs are used, raising concerns for both benchmarking and system deployment. While prior efforts have employed selective input filtering and curated datasets, the impact of input length on perplexity has not been systematically studied from a systems perspective and input length has rarely been treated as a first-class system variable affecting both fairness and efficiency. In this work, we close this gap by introducing LengthBenchmark, a system-conscious evaluation framework that explicitly integrates input length, evaluation protocol design, and system-level costs, evaluating representative LLMs under two scoring protocols (direct accumulation and fixed window sliding) across varying context lengths. Unlike prior work that focuses solely on accuracy-oriented metrics, LengthBenchmark additionally measures latency, memory footprint, and evaluation cost, thereby linking predictive metrics to deployment realities. We further incorporate quantized variants not as a main contribution, but as robustness checks, showing that length-induced biases persist across both full-precision and compressed models. This design disentangles the effects of evaluation logic, quantization, and input length, and demonstrates that length bias is a general phenomenon that undermines fair cross-model comparison. Our analysis yields two key observations: (i) sliding window evaluation consistently inflates performance on short inputs, and (ii) both full-precision and quantized models appear to realise gains as the evaluated segment length grows.

📄 PDF Abstract BibTeX arXiv:2602.04099

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Rethinking GSPO: The Perplexity-Entropy Equivalence

2025-10-27 · Chi Liu arxiv

We provide a new perspective on GSPO's length-normalized importance ratios by establishing their connection to information-theoretic quantities. We show that GSPO's sequence-level weight $s(θ) = (π_θ/π_{θ_{\text{old}}})^…

Mathematical Reasoning

Shortformer: Better Language Modeling using Shorter Inputs

2020-12-31 · ACL 2021 5 · Ofir Press, Noah A. Smith, Mike Lewis

Increasing the input length has been a driver of progress in language modeling with transformers. We identify conditions where shorter inputs are not harmful, and achieve perplexity and efficiency improvements through tw…

Language ModelingLanguage ModellingPositionWord Embeddings

Extending Input Contexts of Language Models through Training on Segmented Sequences

2023-10-23 · Petros Karypis, Julian McAuley, George Karypis

Effectively training language models on long inputs poses many technical challenges. As a cost consideration, languages models are pretrained on a fixed sequence length before being adapted to longer sequences. We explor…

MoQa: Rethinking MoE Quantization with Multi-stage Data-model Distribution Awareness

2025-03-27 · Zihao Zheng, Xiuping Cui, Size Zheng, Maoliang Li 외

With the advances in artificial intelligence, Mix-of-Experts (MoE) has become the main form of Large Language Models (LLMs), and its demand for model compression is increasing. Quantization is an effective method that no…

Language ModelingLanguage ModellingModel CompressionQuantization

Mitigating Forgetting in LLM Fine-Tuning via Low-Perplexity Token Learning

2025-01-24 · Chao-Chung Wu, Zhi Rui Tam, Chieh-Yen Lin, Yun-Nung Chen 외

Maintaining consistent model performance across domains is a fundamental challenge in machine learning. While recent work has explored using LLM-generated data for fine-tuning, its impact on cross-domain generalization r…

Domain Generalization