Bilingual Terminology Extraction Using Neural Word Embeddings on Comparable Corpora
Term and glossary management are vital steps of preparation of every language specialist, and they play a very important role at the stage of education of translation professionals. The growing trend of efficient time management and constant time constraints we may observe in every job sector increases the necessity of the automatic glossary compilation. Many well-performing bilingual AET systems are based on processing parallel data, however, such parallel corpora are not always available for a specific domain or a language pair. Domain-specific, bilingual access to information and its retrieval based on comparable corpora is a very promising area of research that requires a detailed analysis of both available data sources and the possible extraction techniques. This work focuses on domain-specific automatic terminology extraction from comparable corpora for the English – Russian language pair by utilizing neural word embeddings.
Code (0)
등록된 구현이 없습니다.
Tasks
ManagementRetrievalTranslationWord EmbeddingsSimilar Papers 제목 키워드 기반
Bilingual Word Embeddings for Bilingual Terminology Extraction from Specialized Comparable Corpora
Bilingual lexicon extraction from comparable corpora is constrained by the small amount of available data when dealing with specialized domains. This aspect penalizes the performance of distributional-based approaches, w…
Word EmbeddingsLeveraging Meta-Embeddings for Bilingual Lexicon Extraction from Specialized Comparable Corpora
Recent evaluations on bilingual lexicon extraction from specialized comparable corpora have shown contrasted performance while using word embedding models. This can be partially explained by the lack of large specialized…
Information RetrievalMachine TranslationWord Co-occurrence Counts Prediction for Bilingual Terminology Extraction from Comparable Corpora
Towards a unified framework for bilingual terminology extraction of single-word and multi-word terms
Extracting a bilingual terminology for multi-word terms from comparable corpora has not been widely researched. In this work we propose a unified framework for aligning bilingual terms independently of the term lengths. …
Word EmbeddingsImproving Bilingual Terminology Extraction from Comparable Corpora via Multiple Word-Space Models
There is a rich flora of word space models that have proven their efficiency in many different applications including information retrieval (Dumais, 1988), word sense disambiguation (Schutze, 1992), various semantic know…
Information RetrievalRetrievalText CategorizationWord Sense Disambiguation