paper-with-me

홈 › Papers

Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models

2024-05-30 · Zachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion, Matthew L. Leavitt, Mansheej Paul

In this work, we investigate whether small language models can determine high-quality subsets of large-scale text datasets that improve the performance of larger language models. While existing work has shown that pruning based on the perplexity of a larger model can yield high-quality data, we investigate whether smaller models can be used for perplexity-based pruning and how pruning is affected by the domain composition of the data being pruned. We demonstrate that for multiple dataset compositions, perplexity-based pruning of pretraining data can \emph{significantly} improve downstream task performance: pruning based on perplexities computed with a 125 million parameter model improves the average performance on downstream tasks of a 3 billion parameter model by up to 2.04 and achieves up to a $1.45\times$ reduction in pretraining steps to reach commensurate baseline performance. Furthermore, we demonstrate that such perplexity-based data pruning also yields downstream performance gains in the over-trained and data-constrained regimes.

📄 PDF Abstract BibTeX arXiv:2405.20541

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

What Makes My Model Perplexed? A Linguistic Investigation on Neural Language Models Perplexity

2021-06-01 · NAACL (DeeLIO) 2021 6 · Alessio Miaschi, Dominique Brunato, Felice Dell’Orletta, Giulia Venturi

This paper presents an investigation aimed at studying how the linguistic structure of a sentence affects the perplexity of two of the most popular Neural Language Models (NLMs), BERT and GPT-2. We first compare the sent…

Sentence

Large Language Models are Perplexed by some Political Parties

2026-06-04 · Paul Lerner, François Yvon arxiv

Large Language Models (LLMs) are increasingly used, including in political applications, but their political fairness has been little studied. We assess it using perplexity, posing that a fair model should give equal pro…

Is my model perplexed for the right reason? Contrasting LLMs' Benchmark Behavior with Token-Level Perplexity

2026-03-31 · Zoë Prins, Samuele Punzo, Frank Wildenburg, Giovanni Cinà 외 arxiv

Standard evaluations of Large language models (LLMs) focus on task performance, offering limited insight into whether correct behavior reflects appropriate underlying mechanisms and risking confirmation bias. We introduc…

Perplexed by Quality: A Perplexity-based Method for Adult and Harmful Content Detection in Multilingual Heterogeneous Web Data

2022-12-20 · Tim Jansen, Yangling Tong, Victoria Zevallos, Pedro Ortiz Suarez

As demand for large corpora increases with the size of current state-of-the-art language models, using web data as the main part of the pre-training corpus for these models has become a ubiquitous practice. This, in turn…

Language ModellingSmall Language Model

Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar

2025-05-26 · Andrew Gambardella, Takeshi Kojima, Yusuke Iwasawa, Yutaka Matsuo

Typical methods for evaluating the performance of language models evaluate their ability to answer questions accurately. These evaluation metrics are acceptable for determining the extent to which language models can und…

Machine TranslationSentence