paper-with-me

홈 › Papers

Rethinking Model Evaluation as Narrowing the Socio-Technical Gap

2023-06-01 · Q. Vera Liao, Ziang Xiao

The recent development of generative large language models (LLMs) poses new challenges for model evaluation that the research community and industry have been grappling with. While the versatile capabilities of these models ignite much excitement, they also inevitably make a leap toward homogenization: powering a wide range of applications with a single, often referred to as ``general-purpose'', model. In this position paper, we argue that model evaluation practices must take on a critical task to cope with the challenges and responsibilities brought by this homogenization: providing valid assessments for whether and how much human needs in diverse downstream use cases can be satisfied by the given model (\textit{socio-technical gap}). By drawing on lessons about improving research realism from the social sciences, human-computer interaction (HCI), and the interdisciplinary field of explainable AI (XAI), we urge the community to develop evaluation methods based on real-world contexts and human requirements, and embrace diverse evaluation methods with an acknowledgment of trade-offs between realisms and pragmatic costs to conduct the evaluation. By mapping HCI and current NLG evaluation methods, we identify opportunities for evaluation methods for LLMs to narrow the socio-technical gap and pose open questions.

📄 PDF Abstract BibTeX arXiv:2306.03100

Code (0)

등록된 구현이 없습니다.

Tasks

Explainable Artificial Intelligence (XAI)nlg evaluationvalid

Similar Papers 제목 키워드 기반

Rethinking AI Evaluation in Education: The TEACH-AI Framework and Benchmark for Generative AI Assistants

2025-11-28 · Shi Ding, Brian Magerko arxiv

As generative artificial intelligence (AI) continues to transform education, most existing AI evaluations rely primarily on technical performance metrics such as accuracy or task efficiency while overlooking human identi…

Sociotechnical Safety Evaluation of Generative AI Systems

2023-10-18 · Laura Weidinger, Maribeth Rauh, Nahema Marchal, Arianna Manzini 외

Generative AI systems produce a range of risks. To ensure the safety of generative AI systems, these risks must be evaluated. In this paper, we make two main contributions toward establishing such evaluations. First, we …

Policy-as-Prompt: Rethinking Content Moderation in the Age of Large Language Models

2025-02-25 · Konstantina Palla, José Luis Redondo García, Claudia Hauff, Francesco Fabbri 외

Content moderation plays a critical role in shaping safe and inclusive online environments, balancing platform standards, user expectations, and regulatory frameworks. Traditionally, this process involves operationalisin…

Comparing Socio-technical Design Principles with Guidelines for Human-centered AI

2026-07-11 · Thomas Herrmann arxiv

Human-centered AI (HCAI) refers to guidelines or principles that aim on ethi-cally oriented design of systems. We compare HCAI- guidelines with princi-ples of socio-technical systems that emerged in the context of conven…

AI Literacy for All: Adjustable Interdisciplinary Socio-technical Curriculum

2024-09-02 · Sri Yash Tadimalla, Mary Lou Maher

This paper presents a curriculum, "AI Literacy for All," to promote an interdisciplinary understanding of AI, its socio-technical implications, and its practical applications for all levels of education. With the rapid e…

All