paper-with-me

홈 › Papers

Paradigms of AI Evaluation: Mapping Goals, Methodologies and Culture

2025-02-21 · John Burden, Marko Tešić, Lorenzo Pacchiardi, José Hernández-Orallo

Research in AI evaluation has grown increasingly complex and multidisciplinary, attracting researchers with diverse backgrounds and objectives. As a result, divergent evaluation paradigms have emerged, often developing in isolation, adopting conflicting terminologies, and overlooking each other's contributions. This fragmentation has led to insular research trajectories and communication barriers both among different paradigms and with the general public, contributing to unmet expectations for deployed AI systems. To help bridge this insularity, in this paper we survey recent work in the AI evaluation landscape and identify six main paradigms. We characterise major recent contributions within each paradigm across key dimensions related to their goals, methodologies and research cultures. By clarifying the unique combination of questions and approaches associated with each paradigm, we aim to increase awareness of the breadth of current evaluation approaches and foster cross-pollination between different paradigms. We also identify potential gaps in the field to inspire future research directions.

📄 PDF Abstract BibTeX arXiv:2502.15620

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Fragmentation Given a pattern $P,$ that is more complicated than the patterns, we fragment $P$ into simpler patterns such that their exact count is known. In the subgraph GNN proposed earlier,…

Similar Papers 제목 키워드 기반

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

2026-08-18 · Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam 외 arxiv

Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new eviden…

VITAL: A New Dataset for Benchmarking Pluralistic Alignment in Healthcare

2025-02-19 · Anudeex Shetty, Amin Beheshti, Mark Dras, Usman Naseem

Alignment techniques have become central to ensuring that Large Language Models (LLMs) generate outputs consistent with human values. However, existing alignment paradigms often model an averaged or monolithic preference…

BenchmarkingDiversityMultiple-choice

Culture Cartography: Mapping the Landscape of Cultural Knowledge

2025-10-31 · Caleb Ziems, William Held, Jane Yu, Amir Goldberg 외 arxiv

To serve global users safely and productively, LLMs need culture-specific knowledge that might not be learned during pre-training. How do we find such knowledge that is (1) salient to in-group users, but (2) unknown to L…

Culture is Everywhere: A Call for Intentionally Cultural Evaluation

2025-09-01 · Juhyun Oh, Inha Cha, Michael Saxon, Hyunseung Lim 외 arxiv

The prevailing ``trivia-centered paradigm'' for evaluating the cultural alignment of large language models (LLMs) is increasingly inadequate as these models become more advanced and widely deployed. Existing approaches t…

Mapping Patterns for Virtual Knowledge Graphs

2020-12-03 · Diego Calvanese, Avigdor Gal, Davide Lanti, Marco Montali 외

Virtual Knowledge Graphs (VKG) constitute one of the most promising paradigms for integrating and accessing legacy data sources. A critical bottleneck in the integration process involves the definition, validation, and m…

Knowledge GraphsManagement