paper-with-me

홈 › Papers

MUCS@Adap-MT 2020: Low Resource Domain Adaptation for Indic Machine Translation

2020-12-01 · ICON 2020 12 · Asha Hegde, H.l. Shashirekha

Machine Translation (MT) is the task of automatically converting the text in source language to text in target language by preserving the meaning. MT usually require large corpus for training the translation models. Due to scarcity of resources very less attention is given to translating into low resource languages and in particular into Indic languages. In this direction, a shared task called “Adap-MT 2020: Low Resource Domain Adaptation for Indic Machine Translation” is organized to illustrate the capability of general domain MT when translating into Indic languages and low resource domain adaptation of MT systems. In this paper, we, team MUCS, describe a simple word extraction based domain adaptation approach applied to English-Hindi MT only. MT in the proposed model is carried out using Open-NMT - a popular Neural Machine Translation tool. A general domain corpus is built effectively combining the available English-Hindi corpora and removing the duplicate sentences. Further, domain specific corpus is updated by extracting the sentences from generic corpus that contains the words given in the domain specific corpus. The proposed model exhibited satisfactory results for small domain specific AI and CHE corpora provided by the organizers in terms of BLEU score with 1.25 and 2.72 respectively. Further, this methodology is quite generic and can easily be extended to other low resource language pairs as well.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Domain AdaptationMachine TranslationNMTTranslation

Similar Papers 제목 키워드 기반

Exploiting the compressed spectral loss for the learning of the DEMUCS speech enhancement network

2022-11-01 · ROCLING 2022 11 · Chi-En Dai, Qi-Wei Hong, Jeih-weih Hung

This study aims to improve a highly effective speech enhancement technique, DEMUCS, by revising the respective loss function in learning. DEMUCS, developed by Facebook Team, is built on the Wave-UNet and consists of conv…

Speech Enhancement

PDALN: Progressive Domain Adaptation over a Pre-trained Model for Low-Resource Cross-Domain Named Entity Recognition

2021-11-01 · EMNLP 2021 11 · Tao Zhang, Congying Xia, Philip S. Yu, Zhiwei Liu 외

Cross-domain Named Entity Recognition (NER) transfers the NER knowledge from high-resource domains to the low-resource target domain. Due to limited labeled resources and domain shift, cross-domain NER is a challenging t…

Cross-Domain Named Entity RecognitionData AugmentationDomain AdaptationKnowledge Distillation+5

Music Source Separation in the Waveform Domain

2019-11-27 · Alexandre Défossez, Nicolas Usunier, Léon Bottou, Francis Bach

Source separation for music is the task of isolating contributions, or stems, from different instruments recorded individually and arranged together to form a song. Such components include voice, bass, drums and any othe…

Audio GenerationAudio SynthesisData AugmentationMulti-task Audio Source Seperation+2

Hybrid Transformers for Music Source Separation

2022-11-15 · Simon Rouard, Francisco Massa, Alexandre Défossez

A natural question arising in Music Source Separation (MSS) is whether long range contextual information is useful, or whether local acoustic features are sufficient. In other fields, attention based Transformers have sh…

Music Source SeparationSpeech Enhancement

AdapNMT : Neural Machine Translation with Technical Domain Adaptation for Indic Languages

2020-12-01 · ICON 2020 12 · Hema Ala, Dipti Sharma

Adapting new domain is highly challenging task for Neural Machine Translation (NMT). In this paper we show the capability of general domain machine translation when translating into Indic languages (English - Hindi , Eng…

Domain AdaptationMachine TranslationNMTTranslation