paper-with-me

홈 › Papers

Probing the Robustness of Trained Metrics for Conversational Dialogue Systems

2021-11-16 · ACL ARR November 2021 11 · Anonymous

This paper introduces an adversarial method to stress-test trained metrics for the evaluation of conversational dialogue systems. The method leverages Reinforcement Learning to find response strategies that elicit optimal scores from the trained metrics. We apply our method to test recently proposed trained metrics. We find that they all are susceptible to giving high scores to responses generated by rather simple and obviously flawed strategies that our method converges on. For instance, simply copying parts of the conversation context to form a response yields competitive scores or even outperforms responses written by humans.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

reinforcement-learningReinforcement Learning (RL)

Similar Papers 제목 키워드 기반

Probing the Robustness of Trained Metrics for Conversational Dialogue Systems

2022-02-28 · ACL 2022 5 · Jan Deriu, Don Tuggener, Pius von Däniken, Mark Cieliebak

This paper introduces an adversarial method to stress-test trained metrics to evaluate conversational dialogue systems. The method leverages Reinforcement Learning to find response strategies that elicit optimal scores f…

reinforcement-learningReinforcement Learning (RL)

Lightweight Transformers for Conversational AI

2022-07-01 · NAACL (ACL) 2022 7 · Daniel Pressel, Wenshuo Liu, Michael Johnston, Minhua Chen

To understand how training on conversational language impacts performance of pre-trained models on downstream dialogue tasks, we build compact Transformer-based Language Models from scratch on several large corpora of co…

GPUIntent DetectionNatural Language Understanding

Dual Hierarchical Dialogue Policy Learning for Legal Inquisitive Conversational Agents

2026-05-13 · Xubo Lin, Zezhi Deng, Shihao Wang, Grace Hui Yang 외 arxiv

Most existing dialogue systems are user-driven, primarily designed to fulfill user requests. However, in many critical real-world scenarios, a conversational agent must proactively extract information to achieve its own …

Hierarchical Reinforcement Learning

One Battle After Another: Probing LLMs' Limits on Multi-Turn Instruction Following with a Benchmark Evolving Framework

2025-11-05 · Qi Jia, Ye Shen, Xiujie Song, Kaiwei Zhang 외 arxiv

Evaluating LLMs' instruction-following ability in multi-topic dialogues is essential yet challenging. Existing benchmarks are limited to a fixed number of turns, susceptible to saturation and failing to account for users…

Instruction Following

Measuring the Robustness of Reference-Free Dialogue Evaluation Systems

2025-01-12 · Justin Vasselli, Adam Nohejl, Taro Watanabe

Advancements in dialogue systems powered by large language models (LLMs) have outpaced the development of reliable evaluation metrics, particularly for diverse and creative responses. We present a benchmark for evaluatin…

Dialogue EvaluationTAG