paper-with-me

홈 › Papers

Multi2WOZ: A Robust Multilingual Dataset and Conversational Pretraining for Task-Oriented Dialog

2022-05-20 · NAACL 2022 7 · Chia-Chien Hung, Anne Lauscher, Ivan Vulić, Simone Paolo Ponzetto, Goran Glavaš

Research on (multi-domain) task-oriented dialog (TOD) has predominantly focused on the English language, primarily due to the shortage of robust TOD datasets in other languages, preventing the systematic investigation of cross-lingual transfer for this crucial NLP application area. In this work, we introduce Multi2WOZ, a new multilingual multi-domain TOD dataset, derived from the well-established English dataset MultiWOZ, that spans four typologically diverse languages: Chinese, German, Arabic, and Russian. In contrast to concurrent efforts, Multi2WOZ contains gold-standard dialogs in target languages that are directly comparable with development and test portions of the English dataset, enabling reliable and comparative estimates of cross-lingual transfer performance for TOD. We then introduce a new framework for multilingual conversational specialization of pretrained language models (PrLMs) that aims to facilitate cross-lingual transfer for arbitrary downstream TOD tasks. Using such conversational PrLMs specialized for concrete target languages, we systematically benchmark a number of zero-shot and few-shot cross-lingual transfer approaches on two standard TOD tasks: Dialog State Tracking and Response Retrieval. Our experiments show that, in most setups, the best performance entails the combination of (I) conversational specialization in the target language and (ii) few-shot transfer for the concrete TOD task. Most importantly, we show that our conversational specialization in the target language allows for an exceptionally sample-efficient few-shot transfer for downstream TOD tasks.

📄 PDF Abstract BibTeX arXiv:2205.10400

Code (1)

umanlp/multi2woz 공식 구현 pytorch

Tasks

Cross-Lingual Transferdialog state trackingRetrieval

Similar Papers 제목 키워드 기반

Multi2WOZ: A Robust Multilingual Dataset and Conversational Pretraining for Task-Oriented Dialog

2022-01-16 · ACL ARR January 2022 1 · Anonymous

Research on (multi-domain) task-oriented dialog (TOD) has predominantly focused on the English language, primarily due to the shortage of robust TOD datasets in other languages, preventing the systematic investigation of…

Cross-Lingual Transferdialog state trackingRetrieval

Cross-Lingual Pretraining Methods for Spoken Dialog

2021-03-17 · Anonymous

There has been an increasing interest among NLP researchers towards learning generic representations. However, in the field of multilingual spoken dialogue systems, this problem remains overlooked. Indeed most of the pre…

Spoken Dialogue Systems

Addressee and Response Selection for Multilingual Conversation

2018-08-12 · COLING 2018 8 · Motoki Sato, Hiroki Ouch, Yuta Tsuboi

Developing conversational systems that can converse in many languages is an interesting challenge for natural language processing. In this paper, we introduce multilingual addressee and response selection. In this task, …

Transfer Learning

mLongT5: A Multilingual and Efficient Text-To-Text Transformer for Longer Sequences

2023-05-18 · David Uthus, Santiago Ontañón, Joshua Ainslie, Mandy Guo

We present our work on developing a multilingual, efficient text-to-text transformer that is suitable for handling long inputs. This model, called mLongT5, builds upon the architecture of LongT5, while leveraging the mul…

Question Answering

MDAPT: Multilingual Domain Adaptive Pretraining in a Single Model

2021-09-14 · Findings (EMNLP) 2021 11 · Rasmus Kær Jørgensen, Mareike Hartmann, Xiang Dai, Desmond Elliott

Domain adaptive pretraining, i.e. the continued unsupervised pretraining of a language model on domain-specific text, improves the modelling of text for downstream tasks within the domain. Numerous real-world application…

Language ModelingLanguage Modellingnamed-entity-recognitionNamed Entity Recognition+3