paper-with-me

홈 › Papers

Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing

2026-04-01 · Ruchira Dhar, Anders Søgaard arxiv

Recent advances in large language models (LLMs) have prompted a growing body of work that questions the methodology of prevailing evaluation practices. However, many such critiques have already been extensively debated in natural language processing (NLP): a field with a long history of methodological reflection on evaluation. We conduct a scoping review of research on evaluation concerns in NLP and develop a taxonomy, synthesizing recurring positions and trade-offs within each area. We also discuss practical implications of the taxonomy, including a structured checklist to support more deliberate evaluation design and interpretation. By situating contemporary debates within their historical context, this work provides a consolidated reference for reasoning about evaluation practices.

📄 PDF Abstract BibTeX arXiv:2604.25923

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Say It Another Way: A Framework for User-Grounded Paraphrasing

2025-05-06 · Cléa Chataigner, Rebecca Ma, Prakhar Ganesh, Afaf Taïk 외

Small changes in how a prompt is worded can lead to meaningful differences in the behavior of large language models (LLMs), raising concerns about the stability and reliability of their evaluations. While prior work has …

S-Eval: Towards Automated and Comprehensive Safety Evaluation for Large Language Models

2024-05-23 · Xiaohan Yuan, Jinfeng Li, Dongxia Wang, Yuefeng Chen 외

Generative large language models (LLMs) have revolutionized natural language processing with their transformative and emergent capabilities. However, recent evidence indicates that LLMs can produce harmful content that v…

Benchmarking

Revisiting Hotels-50K and Hotel-ID

2022-07-20 · Aarash Feizi, Arantxa Casanova, Adriana Romero-Soriano, Reihaneh Rabbany

In this paper, we propose revisited versions for two recent hotel recognition datasets: Hotels50K and Hotel-ID. The revisited versions provide evaluation setups with different levels of difficulty to better align with th…

Image RetrievalRetrieval

CEB: Compositional Evaluation Benchmark for Fairness in Large Language Models

2024-07-02 · Song Wang, Peng Wang, Tong Zhou, Yushun Dong 외

As Large Language Models (LLMs) are increasingly deployed to handle various natural language processing (NLP) tasks, concerns regarding the potential negative societal impacts of LLM-generated content have also arisen. T…

Fairness

A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations

2025-07-11 · Isar Nejadgholi, Mona Omidyeganeh, Marc-Antoine Drouin, Jonathan Boisvert arxiv

Effective AI governance requires structured approaches for stakeholders to access and verify AI system behavior. With the rise of large language models, Natural Language Explanations (NLEs) are now key to articulating mo…