paper-with-me

Dialogue Evaluation

2개 벤치마크 · 논문 107편 · 이 태스크의 논문 보기 →

Benchmarks

USR-TopicalChat

결과 12개

USR-PersonaChat

결과 10개

Most implemented

Papers

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging

2026-07-11 · Jinglan Gong, Jiefan Lu, Hewei Guo, Kehan Li 외 arxiv

Evaluating large language models (LLMs) as multi-turn conversational partners requires probing capabilities that single-turn benchmarks miss: persona consistency, evolving intent tracking, emotional dynamics, and goal co…

Response GenerationDialogue Evaluation

Cognitive World Models for Process-Level Social Influence Evaluation

2026-06-28 · Minghui Ma, Bin Guo, Han Wang, Mengqi Chen 외 arxiv

Social influence dialogue changes user behavior by altering internal cognitive states. The central evaluation question is whether the user's beliefs, desires, intentions, and emotions measurably change over the course of…

Dialogue Evaluation

Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

2026-06-09 · Atsumoto Ohashi, Neil Zeghidour, Alexandre Défossez, Eugene Kharitonov arxiv

Full-duplex spoken dialogue models can listen and speak simultaneously, making them a promising architecture for natural conversation. However, current models are trained solely with supervised learning through token-lev…

Reinforcement LearningDialogue Evaluation

GRADE: Generalizable Reasoning-Aware Dialogue Evaluation for AI Tutors

2026-05-27 · Parth Bhalerao, Jeromy Chang, David Chou, Oana Ignat arxiv

Evaluating AI tutor responses requires more than factual correctness: tutors must identify mistakes, locate errors, provide guidance, and offer actionable next steps. We present GRADE, a systematic study of open-source m…

Synthetic Data GenerationDialogue Evaluation

Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges

2026-05-13 · Riya Tapwal, Abhishek Kumar, Carsten Maple arxiv

Large language models (LLMs) are increasingly used as automatic judges for summarization and dialogue evaluation. Prior work has documented biases such as position, verbosity, and style preferences, but largely focuses o…

Dialogue Evaluation

Your Students Don't Use LLMs Like You Wish They Did

2026-04-26 · Sebastian Kobler, Matthew Clemson, Angela Sun, Jonathan K. Kummerfeld arxiv

Educational NLP systems are typically evaluated using engagement metrics and satisfaction surveys, which are at best a proxy for meeting pedagogical goals. We introduce six computational metrics for automated evaluation …

Dialogue Evaluation

전체 107편 보기 →