paper-with-me

Papers

Standardizing the Measurement of Text Diversity: A Tool and a Comparative Analysis of Scores

2024-03-01 · Chantal Shaib, Joe Barrow, Jiuding Sun, Alexa F. Siu, Byron C. Wallace, Ani Nenkova

The diversity across outputs generated by large language models shapes the perception of their quality and utility. Prompt leaks, templated answer structure, and canned responses across different interactions are readily noticed by people, but there is no standard score to measure this aspect of model behavior. In this work we empirically investigate diversity scores on English texts. We find that computationally efficient compression algorithms capture information similar to what is measured by slow to compute $n$-gram overlap homogeneity scores. Further, a combination of measures -- compression ratios, self-repetition of long $n$-grams and Self-BLEU and BERTScore -- are sufficient to report, as they have low mutual correlation with each other. The applicability of scores extends beyond analysis of generative models; for example, we highlight applications on instruction-tuning datasets and human-produced texts. We release a diversity score package to facilitate research and invite consistency across reports.

📄 PDF Abstract BibTeX arXiv:2403.00553

Code (1)

cshaib/diversity

Tasks

Diversity

Similar Papers 제목 키워드 기반

emb-diversity: A Tool for Embedding-Based Measurement of Data Diversity

2026-07-22 · Cantao Su, Menan Velayuthan, Esther Ploeger, Dong Nguyen 외 arxiv

There is growing evidence that data diversity is crucial for developing fair and robust NLP models. However, current approaches to measure diversity remain inconsistent and fragmented: While there exist a number of tools…

Artificial Intelligence for Sustainable Urban Biodiversity: A Framework for Monitoring and Conservation

2024-12-28 · Yasmin Rahmati

The rapid expansion of urban areas challenges biodiversity conservation, requiring innovative ecosystem management. This study explores the role of Artificial Intelligence (AI) in urban biodiversity conservation, its app…

Management

Designing Empirical Studies on LLM-Based Code Generation: Towards a Reference Framework

2025-10-04 · Nathalia Nascimento, Everton Guimaraes, Paulo Alencar arxiv

The rise of large language models (LLMs) has introduced transformative potential in automated code generation, addressing a wide range of software engineering challenges. However, empirical evaluation of LLM-based code g…

Code Generation

Measuring Human and AI Values Based on Generative Psychometrics with Large Language Models

2024-09-18 · Haoran Ye, Yuhang Xie, Yuanyi Ren, Hanjun Fang 외

Human values and their measurement are long-standing interdisciplinary inquiry. Recent advances in AI have sparked renewed interest in this area, with large language models (LLMs) emerging as both tools and subjects of v…

Out-of-the-Box and into the Ditch? Multilingual Evaluation of Generic Text Extraction Tools

2020-05-01 · LREC 2020 5 · Adrien Barbaresi, Ga{\"e}l Lejeune

This article examines extraction methods designed to retain the main text content of web pages and discusses how the extraction could be oriented and evaluated: can and should it be as generic as possible to ensure oppor…

Diversity