paper-with-me

홈 › Papers

From Masked Language Modeling to Translation: Non-English Auxiliary Tasks Improve Zero-shot Spoken Language Understanding

2021-05-15 · NAACL 2021 4 · Rob van der Goot, Ibrahim Sharaf, Aizhan Imankulova, Ahmet Üstün, Marija Stepanović, Alan Ramponi, Siti Oryza Khairunnisa, Mamoru Komachi, Barbara Plank

The lack of publicly available evaluation data for low-resource languages limits progress in Spoken Language Understanding (SLU). As key tasks like intent classification and slot filling require abundant training data, it is desirable to reuse existing data in high-resource languages to develop models for low-resource scenarios. We introduce xSID, a new benchmark for cross-lingual Slot and Intent Detection in 13 languages from 6 language families, including a very low-resource dialect. To tackle the challenge, we propose a joint learning approach, with English SLU training data and non-English auxiliary tasks from raw text, syntax and translation for transfer. We study two setups which differ by type and language coverage of the pre-trained embeddings. Our results show that jointly learning the main tasks with masked language modeling is effective for slots, while machine translation transfer works best for intent classification.

📄 PDF Abstract BibTeX arXiv:2105.07316

Code (2)

https://bitbucket.org/robvanderg/xsid 공식 구현 pytorch
Kaleidophon/deep-significance 공식 구현 tf

Tasks

intent-classificationIntent ClassificationIntent DetectionLanguage ModelingLanguage ModellingMachine TranslationMasked Language ModelingSlot FillingSpoken Language UnderstandingTranslation

Similar Papers 제목 키워드 기반

Learning to Reuse Translations: Guiding Neural Machine Translation with Examples

2019-11-25 · Qian Cao, Shaohui Kuang, Deyi Xiong

In this paper, we study the problem of enabling neural machine translation (NMT) to reuse previous translations from similar examples in target prediction. Distinguishing reusable translations from noisy segments and lea…

DecoderMachine TranslationNMTTranslation

UC2: Universal Cross-lingual Cross-modal Vision-and-Language Pre-training

2021-04-01 · CVPR 2021 1 · Mingyang Zhou, Luowei Zhou, Shuohang Wang, Yu Cheng 외

Vision-and-language pre-training has achieved impressive success in learning multimodal representations between vision and language. To generalize this success to non-English languages, we introduce UC2, the first machin…

Image-text matchingImage-text RetrievalLanguage ModelingLanguage Modelling+10

Semantic Alignment across Ancient Egyptian Language Stages via Normalization-Aware Multitask Learning

2026-03-25 · He Huang arxiv

We study word-level semantic alignment across four historical stages of Ancient Egyptian. These stages differ in script and orthography, and parallel data are scarce. We jointly train a compact encoder-decoder model with…

Part-Of-Speech Tagging

Women Are Beautiful, Men Are Leaders: Gender Stereotypes in Machine Translation and Language Modeling

2023-11-30 · Matúš Pikuliak, Andrea Hrckova, Stefan Oresko, Marián Šimko

We present GEST -- a new manually created dataset designed to measure gender-stereotypical reasoning in language models and machine translation systems. GEST contains samples for 16 gender stereotypes about men and women…

Language ModelingLanguage ModellingMachine TranslationTranslation

Multilingual Speech-to-Speech Translation into Multiple Target Languages

2023-07-17 · Hongyu Gong, Ning Dong, Sravya Popuri, Vedanuj Goswami 외

Speech-to-speech translation (S2ST) enables spoken communication between people talking in different languages. Despite a few studies on multilingual S2ST, their focus is the multilinguality on the source side, i.e., the…

Language IdentificationSpeech-to-Speech TranslationTranslation