paper-with-me

Papers

Linguistically-driven Multi-task Pre-training for Low-resource Neural Machine Translation

2022-01-20 · Zhuoyuan Mao, Chenhui Chu, Sadao Kurohashi

In the present study, we propose novel sequence-to-sequence pre-training objectives for low-resource machine translation (NMT): Japanese-specific sequence to sequence (JASS) for language pairs involving Japanese as the source or target language, and English-specific sequence to sequence (ENSS) for language pairs involving English. JASS focuses on masking and reordering Japanese linguistic units known as bunsetsu, whereas ENSS is proposed based on phrase structure masking and reordering tasks. Experiments on ASPEC Japanese--English & Japanese--Chinese, Wikipedia Japanese--Chinese, News English--Korean corpora demonstrate that JASS and ENSS outperform MASS and other existing language-agnostic pre-training methods by up to +2.9 BLEU points for the Japanese--English tasks, up to +7.0 BLEU points for the Japanese--Chinese tasks and up to +1.3 BLEU points for English--Korean tasks. Empirical analysis, which focuses on the relationship between individual parts in JASS and ENSS, reveals the complementary nature of the subtasks of JASS and ENSS. Adequacy evaluation using LASER, human evaluation, and case studies reveals that our proposed methods significantly outperform pre-training methods without injected linguistic knowledge and they have a larger positive impact on the adequacy as compared to the fluency. We release codes here: https://github.com/Mao-KU/JASS/tree/master/linguistically-driven-pretraining.

📄 PDF Abstract BibTeX arXiv:2201.08070

Code (1)

Mao-KU/JASS 공식 구현

Tasks

Low Resource Neural Machine TranslationLow-Resource Neural Machine TranslationMachine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

Uncovering Code-Mixed Challenges: A Framework for Linguistically Driven Question Generation and Neural Based Question Answering

2018-10-01 · CONLL 2018 10 · Deepak Gupta, Pabitra Lenka, Asif Ekbal, Pushpak Bhattacharyya

Existing research on question answering (QA) and comprehension reading (RC) are mainly focused on the resource-rich language like English. In recent times, the rapid growth of multi-lingual web content has posed several …

Question AnsweringQuestion GenerationQuestion-Generation

Improving Natural Language Inference in Arabic using Transformer Models and Linguistically Informed Pre-Training

2023-07-27 · Mohammad Majd Saad Al Deen, Maren Pielka, Jörn Hees, Bouthaina Soulef Abdou 외

This paper addresses the classification of Arabic text data in the field of Natural Language Processing (NLP), with a particular focus on Natural Language Inference (NLI) and Contradiction Detection (CD). Arabic is consi…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)Natural Language Inference+1

NSL-MT: Linguistically Informed Negative Samples for Efficient Machine Translation in Low-Resource Languages

2025-11-12 · Mamadou K. Keita, Christopher Homan, Huy Le arxiv

We introduce negative space learning machine translation (NSL-MT), a training method for underresourced languages, that augments limited parallel data with synthetically generated violations of the target language's gram…

Machine Translation

Unsupervised Linguistically-Driven Reliable Dependency Parses Detection and Self-Training for Adaptation to the Biomedical Domain

2013-08-01 · WS 2013 8 · Felice Dell{'}Orletta, Giulia Venturi, Simonetta Montemagni
Domain AdaptationPart-Of-Speech Tagging

Linguistically-Driven Strategy for Concept Prerequisites Learning on Italian

2019-08-01 · WS 2019 8 · Alessio Miaschi, Chiara Alzetta, Franco Alberto Cardillo, Felice Dell{'}Orletta

We present a new concept prerequisite learning method for Learning Object (LO) ordering that exploits only linguistic features extracted from textual educational resources. The method was tested in a cross- and in- domai…