paper-with-me

홈 › Papers

Low-Resource Unsupervised NMT: Diagnosing the Problem and Providing a Linguistically Motivated Solution

2020-11-01 · EAMT 2020 11 · Lukas Edman, Antonio Toral, Gertjan van Noord

Unsupervised Machine Translation has been advancing our ability to translate without parallel data, but state-of-the-art methods assume an abundance of monolingual data. This paper investigates the scenario where monolingual data is limited as well, finding that current unsupervised methods suffer in performance under this stricter setting. We find that the performance loss originates from the poor quality of the pretrained monolingual embeddings, and we offer a potential solution: dependency-based word embeddings. These embeddings result in a complementary word representation which offers a boost in performance of around 1.5 BLEU points compared to standard word2vec when monolingual data is limited to 1 million sentences per language. We also find that the inclusion of sub-word information is crucial to improving the quality of the embeddings.

📄 PDF Abstract BibTeX

Code (1)

leukas/lrumt 공식 구현 pytorch

Tasks

Machine TranslationNMTTranslationUnsupervised Machine TranslationWord Embeddings

Similar Papers 제목 키워드 기반

Exploiting Cross-Lingual Knowledge in Unsupervised Acoustic Modeling for Low-Resource Languages

2020-07-29 · Siyuan Feng

(Short version of Abstract) This thesis describes an investigation on unsupervised acoustic modeling (UAM) for automatic speech recognition (ASR) in the zero-resource scenario, where only untranscribed speech data is ass…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language Acquisitionspeech-recognition+1

The SETimes.HR Linguistically Annotated Corpus of Croatian

2014-05-01 · LREC 2014 5 · {\v{Z}}eljko Agi{\'c}, Nikola Ljube{\v{s}}i{\'c}

We present SETimes.HR ― the first linguistically annotated corpus of Croatian that is freely available for all purposes. The corpus is built on top of the SETimes parallel corpus of nine Southeast European languages an…

AllBoundary DetectionDependency ParsingLemmatization+4

A Little Linguistics Goes a Long Way: Unsupervised Segmentation with Limited Language Specific Guidance

2019-08-01 · WS 2019 8 · Alex Erdmann, er, Salam Khalifa, Mai Oudah 외

We present de-lexical segmentation, a linguistically motivated alternative to greedy or other unsupervised methods, requiring only minimal language specific input. Our technique involves creating a small grammar of close…

When Less Is More? Diagnosing ASR Predictions in Sardinian via Layer-Wise Decoding

2026-02-10 · Domenico De Cristofaro, Alessandro Vietti, Marianne Pouplier, Aleese Block arxiv

Recent studies have shown that intermediate layers in multilingual speech models often encode more phonetically accurate representations than the final output layer. In this work, we apply a layer-wise decoding strategy …

Unsupervised Linguistically-Driven Reliable Dependency Parses Detection and Self-Training for Adaptation to the Biomedical Domain

2013-08-01 · WS 2013 8 · Felice Dell{'}Orletta, Giulia Venturi, Simonetta Montemagni
Domain AdaptationPart-Of-Speech Tagging