paper-with-me

Papers

Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents

2026-08-06 · Noam Koren, Roy Bar-Haim, Abigail Goldsteen arxiv

Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplistic scenarios, or limited policy coverage, leading to unreliable evaluations. We introduce a reference-free framework that uses LLM judges to assess benchmark consistency, complexity, and policy coverage, while providing actionable diagnostics of weaknesses. We validate the framework by demonstrating agreement with independent human annotations and by evaluating benchmarks generated by LLMs of varying capabilities, as well as benchmarks subjected to controlled quality-degrading perturbations. Across domains and judge models, the proposed metrics consistently distinguish between benchmark quality levels. We further demonstrate the framework's applicability to manually curated benchmarks. Our framework offers a practical approach for evaluating synthetic and manually curated conversational-agent benchmarks.

📄 PDF Abstract BibTeX arXiv:2608.06329

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents

2026-01-07 · Yuanchen Bei, Tianxin Wei, Xuying Ning, Yanjun Zhao 외 arxiv

Long-term memory is a critical capability for multimodal large language model (MLLM) agents, particularly in conversational settings where information accumulates and evolves over time. However, existing benchmarks eithe…

Test-time Adaptation

MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks

2026-05-20 · Junhao Ruan, Abudukeyumu Abudula, Bei Li, Yongjing Yin 외 arxiv

Accurate evaluation of conversational retrieval is pivotal for advancing Retrieval-Augmented Generation (RAG) systems. However, existing conversational retrieval benchmarks suffer from costly, sparse human annotation or …

M$^3$Exam: Benchmarking Multimodal Memory for Realistic User-Agent Interactions

2026-06-05 · Zhengjun Huang, Wenxuan Liu, Zhoujin Tian, Wei Chen 외 arxiv

Language agents are increasingly deployed over accumulating multimodal information, yet existing benchmarks assume a human-human form with sparse visuals and straightforward content, evaluating neither reasoning over aut…

GiCCS: A German in-Context Conversational Similarity Benchmark

2022-12-16 · Proceedings of the 2nd Workshop on Natural Language Generation, Evaluation, and Metrics (GEM) 2022 12 · Shima Asaadi, Zahra Kolagar, Alina Liebel, Alessandra Zarcone

The Semantic textual similarity (STS) task is commonly used to evaluate the semantic representations that language models (LMs) learn from texts, under the assumption that good-quality representations will yield accurate…

BenchmarkingSemantic Textual SimilaritySTS

Dynamic benchmarking framework for LLM-based conversational data capture

2025-02-04 · Pietro Alessandro Aluffi, Patrick Zietkiewicz, Marya Bazzi, Matt Arderne 외

The rapid evolution of large language models (LLMs) has transformed conversational agents, enabling complex human-machine interactions. However, evaluation frameworks often focus on single tasks, failing to capture the d…

Benchmarking