paper-with-me

홈 › Papers

Efficient Data Selection for Bilingual Terminology Extraction from Comparable Corpora

2016-12-01 · COLING 2016 12 · Amir Hazem, Emmanuel Morin

Comparable corpora are the main alternative to the use of parallel corpora to extract bilingual lexicons. Although it is easier to build comparable corpora, specialized comparable corpora are often of modest size in comparison with corpora issued from the general domain. Consequently, the observations of word co-occurrences which are the basis of context-based methods are unreliable. We propose in this article to improve word co-occurrences of specialized comparable corpora and thus context representation by using general-domain data. This idea, which has been already used in machine translation task for more than a decade, is not straightforward for the task of bilingual lexicon extraction from specific-domain comparable corpora. We go against the mainstream of this task where many studies support the idea that adding out-of-domain documents decreases the quality of lexicons. Our empirical evaluation shows the advantages of this approach which induces a significant gain in the accuracy of extracted lexicons.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine TranslationTopic ModelsTranslationWord Embeddings

Similar Papers 제목 키워드 기반

Leveraging Meta-Embeddings for Bilingual Lexicon Extraction from Specialized Comparable Corpora

2018-08-01 · COLING 2018 8 · Amir Hazem, Emmanuel Morin

Recent evaluations on bilingual lexicon extraction from specialized comparable corpora have shown contrasted performance while using word embedding models. This can be partially explained by the lack of large specialized…

Information RetrievalMachine Translation

Word Co-occurrence Counts Prediction for Bilingual Terminology Extraction from Comparable Corpora

2013-10-01 · IJCNLP 2013 10 · Amir Hazem, Emmanuel Morin

Bilingual Word Embeddings for Bilingual Terminology Extraction from Specialized Comparable Corpora

2017-11-01 · IJCNLP 2017 11 · Amir Hazem, Emmanuel Morin

Bilingual lexicon extraction from comparable corpora is constrained by the small amount of available data when dealing with specialized domains. This aspect penalizes the performance of distributional-based approaches, w…

Word Embeddings

Bilingual Terminology Extraction Using Neural Word Embeddings on Comparable Corpora

2021-09-01 · RANLP 2021 9 · Darya Filippova, Burcu Can, Gloria Corpas Pastor

Term and glossary management are vital steps of preparation of every language specialist, and they play a very important role at the stage of education of translation professionals. The growing trend of efficient time ma…

ManagementRetrievalTranslationWord Embeddings

Bilingual Terminology Extraction from Comparable E-Commerce Corpora

2021-04-15 · Hao Jia, Shuqin Gu, Yuqi Zhang, Xiangyu Duan

Bilingual terminologies are important machine translation resources in the field of e-commerce, which are usually either manually translated or automatically extracted from parallel data. The human translation is costly …

Machine TranslationSentenceTranslation