paper-with-me

홈 › Papers

The SADID Evaluation Datasets for Low-Resource Spoken Language Machine Translation of Arabic Dialects

2020-12-01 · COLING 2020 8 · Wael Abid

Low-resource Machine Translation recently gained a lot of popularity, and for certain languages, it has made great strides. However, it is still difficult to track progress in other languages for which there is no publicly available evaluation data. In this paper, we introduce benchmark datasets for Arabic and its dialects. We describe our design process and motivations and analyze the datasets to understand their resulting properties. Numerous successful attempts use large monolingual corpora to augment low-resource pairs. We try to approach augmentation differently and investigate whether it is possible to improve MT models without any external sources of data. We accomplish this by bootstrapping existing parallel sentences and complement this with multilingual training to achieve strong baselines.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTranslation

Similar Papers 제목 키워드 기반

Language Agnostic Data-Driven Inverse Text Normalization

2023-01-20 · Szu-Jui Chen, Debjyoti Paul, Yutong Pang, Peng Su 외

With the emergence of automatic speech recognition (ASR) models, converting the spoken form text (from ASR) to the written form is in urgent need. This inverse text normalization (ITN) problem attracts the attention of r…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Data AugmentationForm+3

SART - Similarity, Analogies, and Relatedness for Tatar Language: New Benchmark Datasets for Word Embeddings Evaluation

2019-03-31 · Albina Khusainova, Adil Khan, Adín Ramírez Rivera

There is a huge imbalance between languages currently spoken and corresponding resources to study them. Most of the attention naturally goes to the "big" languages: those which have the largest presence in terms of media…

Embeddings EvaluationLanguage ModelingLanguage ModellingWord Embeddings

Multilingual and Cross-Lingual Intent Detection from Spoken Data

2021-04-17 · EMNLP 2021 11 · Daniela Gerz, Pei-Hao Su, Razvan Kusztos, Avishek Mondal 외

We present a systematic study on multilingual and cross-lingual intent detection from spoken data. The study leverages a new resource put forth in this work, termed MInDS-14, a first training and evaluation resource for …

Few-Shot LearningIntent DetectionMachine TranslationSentence+3

The SI TEDx-UM speech database: a new Slovenian Spoken Language Resource

2016-05-01 · LREC 2016 5 · Andrej {\v{Z}}gank, Mirjam Sepesy Mau{\v{c}}ec, Darinka Verdonik

This paper presents a new Slovenian spoken language resource built from TEDx Talks. The speech database contains 242 talks in total duration of 54 hours. The annotation and transcription of acquired spoken material was g…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+2

SLUE: New Benchmark Tasks for Spoken Language Understanding Evaluation on Natural Speech

2021-11-19 · Suwon Shon, Ankita Pasad, Felix Wu, Pablo Brusco 외

Progress in speech processing has been facilitated by shared datasets and benchmarks. Historically these have focused on automatic speech recognition (ASR), speaker identification, or other lower-level tasks. Interest ha…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)named-entity-recognitionNamed Entity Recognition+6