paper-with-me

Papers

Deconstructing NLG Evaluation: Evaluation Practices, Assumptions, and Their Implications

2022-05-13 · NAACL 2022 7 · Kaitlyn Zhou, Su Lin Blodgett, Adam Trischler, Hal Daumé III, Kaheer Suleman, Alexandra Olteanu

There are many ways to express similar things in text, which makes evaluating natural language generation (NLG) systems difficult. Compounding this difficulty is the need to assess varying quality criteria depending on the deployment setting. While the landscape of NLG evaluation has been well-mapped, practitioners' goals, assumptions, and constraints -- which inform decisions about what, when, and how to evaluate -- are often partially or implicitly stated, or not stated at all. Combining a formative semi-structured interview study of NLG practitioners (N=18) with a survey study of a broader sample of practitioners (N=61), we surface goals, community practices, assumptions, and constraints that shape NLG evaluations, examining their implications and how they embody ethical considerations.

📄 PDF Abstract BibTeX arXiv:2205.06828

Code (0)

등록된 구현이 없습니다.

Tasks

nlg evaluationText Generation

Similar Papers 제목 키워드 기반

Neither Valid nor Reliable? Investigating the Use of LLMs as Judges

2025-08-25 · Khaoula Chehbouni, Mohammed Haddou, Jackie Chi Kit Cheung, Golnoosh Farnadi arxiv

Evaluating natural language generation (NLG) systems remains a core challenge of natural language processing (NLP), further complicated by the rise of large language models (LLMs) that aims to be general-purpose. Recentl…

Text Summarization

Position: AI Evaluations Should be Grounded on a Theory of Capability

2025-09-23 · Nathanael Jo, Ashia Wilson arxiv

Evaluations of generative models are now ubiquitous, and their outcomes critically shape public and scientific expectations of AI's capabilities. Yet skepticism about their reliability continues to grow. How can we know …

Lost in Speech: Benchmarking, Evaluation, and Parsing of Spoken Bilingual Conversational Language Beyond Standard UD Assumptions

2026-02-06 · Nemika Tyagi, Olga Kellert, Holly Hendrix, Nelvin Licona-Guevara 외 arxiv

Spoken bilingual conversations pose substantial challenges for syntactic parsing because they often include disfluencies and discourse-driven structures that complicate dependency parsing under standard Universal Depende…

Dependency Parsing

Culture is Everywhere: A Call for Intentionally Cultural Evaluation

2025-09-01 · Juhyun Oh, Inha Cha, Michael Saxon, Hyunseung Lim 외 arxiv

The prevailing ``trivia-centered paradigm'' for evaluating the cultural alignment of large language models (LLMs) is increasingly inadequate as these models become more advanced and widely deployed. Existing approaches t…

Interactive Evaluation Requires a Design Science

2026-05-18 · Keyang Xuan, Peiyang Song, Pan Lu, Pengrui Han 외 arxiv

AI evaluation is undergoing a structural change. Large language models (LLMs) are increasingly deployed as systems that act over time through tools, environments, users, and other agents, while many evaluation practices …