Dialogue Evaluation
2개 벤치마크 · 논문 107편 · 이 태스크의 논문 보기 →
Benchmarks
USR-TopicalChat
USR-PersonaChat
Most implemented
Adversarial Learning for Neural Dialogue Generation
Don't Forget Your ABC's: Evaluating the State-of-the-Art in Chat-Oriented Dialogue Systems
FineD-Eval: Fine-grained Automatic Dialogue-Level Evaluation
Automatic Evaluation and Moderation of Open-domain Dialogue Systems
Unsupervised Evaluation of Interactive Dialog with DialoGPT
Papers
Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging
Evaluating large language models (LLMs) as multi-turn conversational partners requires probing capabilities that single-turn benchmarks miss: persona consistency, evolving intent tracking, emotional dynamics, and goal co…
Response GenerationDialogue EvaluationCognitive World Models for Process-Level Social Influence Evaluation
Social influence dialogue changes user behavior by altering internal cognitive states. The central evaluation question is whether the user's beliefs, desires, intentions, and emotions measurably change over the course of…
Dialogue EvaluationMulti-Faceted Interactivity Alignment in Full-Duplex Speech Models
Full-duplex spoken dialogue models can listen and speak simultaneously, making them a promising architecture for natural conversation. However, current models are trained solely with supervised learning through token-lev…
Reinforcement LearningDialogue EvaluationGRADE: Generalizable Reasoning-Aware Dialogue Evaluation for AI Tutors
Evaluating AI tutor responses requires more than factual correctness: tutors must identify mistakes, locate errors, provide guidance, and offer actionable next steps. We present GRADE, a systematic study of open-source m…
Synthetic Data GenerationDialogue EvaluationFaithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges
Large language models (LLMs) are increasingly used as automatic judges for summarization and dialogue evaluation. Prior work has documented biases such as position, verbosity, and style preferences, but largely focuses o…
Dialogue EvaluationYour Students Don't Use LLMs Like You Wish They Did
Educational NLP systems are typically evaluated using engagement metrics and satisfaction surveys, which are at best a proxy for meeting pedagogical goals. We introduce six computational metrics for automated evaluation …
Dialogue Evaluation