paper-with-me

Papers

Building Comparable Corpora for Assessing Multi-Word Term Alignment

2022-06-01 · LREC 2022 6 · Omar Adjali, Emmanuel Morin, Pierre Zweigenbaum

Recent work has demonstrated the importance of dealing with Multi-Word Terms (MWTs) in several Natural Language Processing applications. In particular, MWTs pose serious challenges for alignment and machine translation systems because of their syntactic and semantic properties. Thus, developing algorithms that handle MWTs is becoming essential for many NLP tasks. However, the availability of bilingual and more generally multi-lingual resources is limited, especially for low-resourced languages and in specialized domains. In this paper, we propose an approach for building comparable corpora and bilingual term dictionaries that help evaluate bilingual term alignment in comparable corpora. To that aim, we exploit parallel corpora to perform automatic bilingual MWT extraction and comparable corpus construction. Parallel information helps to align bilingual MWTs and makes it easier to build comparable specialized sub-corpora. Experimental validation on an existing dataset and on manually annotated data shows the interest of the proposed methodology.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Machine Translation

Similar Papers 제목 키워드 기반

Assessing Back-Translation as a Corpus Generation Strategy for non-English Tasks: A Study in Reading Comprehension and Word Sense Disambiguation

2019-08-01 · WS 2019 8 · Fabricio Monsalve, Kervy Rivas Rojas, Marco Antonio Sobrevilla Cabezudo, Arturo Oncevay

Corpora curated by experts have sustained Natural Language Processing mainly in English, but the expensiveness of corpora creation is a barrier for the development in further languages. Thus, we propose a corpus generati…

Machine TranslationReading ComprehensionTranslationWord Sense Disambiguation

Assessing the Corpus Size vs. Similarity Trade-off for Word Embeddings in Clinical NLP

2016-12-01 · WS 2016 12 · Kirk Roberts

The proliferation of deep learning methods in natural language processing (NLP) and the large amounts of data they often require stands in stark contrast to the relatively data-poor clinical NLP domain. In particular, la…

Deep LearningWord Embeddings

Overview of the Fourth BUCC Shared Task: Bilingual Dictionary Induction from Comparable Corpora

2020-05-01 · LREC 2020 5 · Reinhard Rapp, Pierre Zweigenbaum, Serge Sharoff

The shared task of the 13th Workshop on Building and Using Comparable Corpora was devoted to the induction of bilingual dictionaries from comparable rather than parallel corpora. In this task, for a number of language pa…

Bilingual Word Embeddings for Bilingual Terminology Extraction from Specialized Comparable Corpora

2017-11-01 · IJCNLP 2017 11 · Amir Hazem, Emmanuel Morin

Bilingual lexicon extraction from comparable corpora is constrained by the small amount of available data when dealing with specialized domains. This aspect penalizes the performance of distributional-based approaches, w…

Word Embeddings

Harvesting comparable corpora and mining them for equivalent bilingual sentences using statistical classification and analogy- based heuristics

2015-11-18 · Krzysztof Wołk, Emilia Rejmund, Krzysztof Marasek

Parallel sentences are a relatively scarce but extremely useful resource for many applications including cross-lingual retrieval and statistical machine translation. This research explores our new methodologies for minin…

General ClassificationMachine TranslationRetrievalTranslation