paper-with-me

Papers

RADDLE: An Evaluation Benchmark and Analysis Platform for Robust Task-oriented Dialog Systems

2020-12-29 · ACL 2021 5 · Baolin Peng, Chunyuan Li, Zhu Zhang, Chenguang Zhu, Jinchao Li, Jianfeng Gao

For task-oriented dialog systems to be maximally useful, it must be able to process conversations in a way that is (1) generalizable with a small number of training examples for new task domains, and (2) robust to user input in various styles, modalities or domains. In pursuit of these goals, we introduce the RADDLE benchmark, a collection of corpora and tools for evaluating the performance of models across a diverse set of domains. By including tasks with limited training data, RADDLE is designed to favor and encourage models with a strong generalization ability. RADDLE also includes a diagnostic checklist that facilitates detailed robustness analysis in aspects such as language variations, speech errors, unseen entities, and out-of-domain utterances. We evaluate recent state-of-the-art systems based on pre-training and fine-tuning, and find that grounded pre-training on heterogeneous dialog corpora performs better than training a separate model per domain. Overall, existing models are less than satisfactory in robustness evaluation, which suggests opportunities for future improvement.

📄 PDF Abstract BibTeX arXiv:2012.14666

Code (0)

등록된 구현이 없습니다.

Tasks

Diagnostic

Similar Papers 제목 키워드 기반

Supervised machine learning classification for short straddles on the S&P500

2022-04-26 · Alexander Brunhuemer, Lukas Larcher, Philipp Seidl, Sascha Desmettre 외

In this working paper we present our current progress in the training of machine learning models to execute short option strategies on the S&P500. As a first step, this paper is breaking this problem down to a supervised…

BIG-bench Machine Learning

Using linear initialisation to improve speed of convergence and fully-trained error in Autoencoders

2023-11-17 · Marcel Marais, Mate Hartstein, George Cevora

Good weight initialisation is an important step in successful training of Artificial Neural Networks. Over time a number of improvements have been proposed to this process. In this paper we introduce a novel weight initi…

Active Learning for Level Set Estimation Using Randomized Straddle Algorithms

2024-08-06 · Yu Inatsu, Shion Takeno, Kentaro Kutsukake, Ichiro Takeuchi

Level set estimation (LSE), the problem of identifying the set of input points where a function takes value above (or below) a given threshold, is important in practical applications. When the function is expensive-to-ev…

Active Learning

Modular Anthropomorphic Hand Design via Multi-Parameter Finger Benchmarking and Selection

2026-06-10 · Yu Zhang, Huijiang Wang, Josie Hughes arxiv

Designing anthropomorphic dexterous robotic hands remains challenging as the design space straddles morphology, actuation, and sensing properties, and performance metrics span both task-dependent and task-agnostic. Exist…

Reimagining GNN Explanations with ideas from Tabular Data

2021-06-23 · Anjali Singh, Shamanth R Nayak K, Balaji Ganesan

Explainability techniques for Graph Neural Networks still have a long way to go compared to explanations available for both neural and decision decision tree-based models trained on tabular data. Using a task that stradd…