paper-with-me

홈 › Papers

NC-Bench: An LLM Benchmark for Evaluating Conversational Competence

2026-01-10 · Robert J. Moore, Sungeun An, Farhan Ahmed, Jay Pankaj Gala arxiv

The Natural Conversation Benchmark (NC-Bench) introduces a new approach to evaluating the general conversational competence of large language models (LLMs). Unlike prior benchmarks that focus on the content of model behavior, NC-Bench focuses on the form and structure of natural conversation. Grounded in the IBM Natural Conversation Framework (NCF), NC-Bench comprises three distinct sets: (1) the basic set evaluates fundamental sequence management practices, such as answering inquiries, repairing responses, and closing conversational pairs; (2) the retrieval-augmented generation (RAG) set applies the same sequence management patterns as the first set but incorporates information-seeking via RAG; (3) the complex request set extends to requests involving more intricate sequence management patterns. Each set tests a model's ability to produce contextually appropriate conversational actions in response to characteristic interaction patterns. Initial evaluations across six open-source models and 14 interaction patterns show that models perform well on basic answering tasks, struggle more with repair tasks (especially repeat), have mixed performance on closing sequences, and find complex multi-turn requests most challenging. By operationalizing fundamental principles of human conversation, NC-Bench provides a lightweight, extensible, and theory-grounded framework for assessing and improving the conversational abilities of LLMs beyond topical or task-specific benchmarks.

📄 PDF Abstract BibTeX arXiv:2601.06426

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

RecToM: A Benchmark for Evaluating Machine Theory of Mind in LLM-based Conversational Recommender Systems

2025-11-27 · Mengfan Li, Xuanhua Shi, Yang Deng arxiv

Large Language models are revolutionizing the conversational recommender systems through their impressive capabilities in instruction comprehension, reasoning, and human interaction. A core factor underlying effective re…

Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents

2026-08-06 · Noam Koren, Roy Bar-Haim, Abigail Goldsteen arxiv

Task-oriented conversational agents are evaluated using curated or automatically generated benchmarks, yet benchmark quality is rarely assessed. Poor benchmarks may contain inconsistent tasks, simplistic scenarios, or li…

Pardon? Evaluating Conversational Repair in Large Audio-Language Models

2026-01-19 · Shuanghong Huang, Jinlei Xu, Youchao Zhou, Yanghao Zhou 외 arxiv

Large Audio-Language Models (LALMs) have demonstrated strong performance in spoken question answering (QA), with existing evaluations primarily focusing on answer accuracy and robustness to acoustic perturbations. Howeve…

Question Answering

CPG-EVAL: A Multi-Tiered Benchmark for Evaluating the Chinese Pedagogical Grammar Competence of Large Language Models

2025-04-17 · Dong Wang

Purpose: The rapid emergence of large language models (LLMs) such as ChatGPT has significantly impacted foreign language education, yet their pedagogical grammar competence remains under-assessed. This paper introduces C…

CCBENCH: Assessing LLM Cultural Competence via Implicitly Signaled Norms using Health Queries

2026-06-08 · Vasudha Varadarajan, Akhila Yerukola, Mona T. Diab, Maarten Sap arxiv

To interact with users fairly and without stereotyping, AI models must display cultural competency, i.e., the ability to infer and adapt to a user's implicitly signaled cultural values, rather than relying on static demo…