Sub-corpora Sampling with an Application to Bilingual Lexicon Extraction
Code (0)
등록된 구현이 없습니다.
Tasks
Data AugmentationSimilar Papers 제목 키워드 기반
Discovering Bilingual Lexicons in Polyglot Word Embeddings
Bilingual lexicons and phrase tables are critical resources for modern Machine Translation systems. Although recent results show that without any seed lexicon or parallel data, highly accurate bilingual lexicons can be l…
Machine TranslationTranslationWord EmbeddingsAdaptive Dictionary for Bilingual Lexicon Extraction from Comparable Corpora
One of the main resources used for the task of bilingual lexicon extraction from comparable corpora is : the bilingual dictionary, which is considered as a bridge between two languages. However, no particular attention h…
Information RetrievalLeveraging Meta-Embeddings for Bilingual Lexicon Extraction from Specialized Comparable Corpora
Recent evaluations on bilingual lexicon extraction from specialized comparable corpora have shown contrasted performance while using word embedding models. This can be partially explained by the lack of large specialized…
Information RetrievalMachine TranslationEfficient Data Selection for Bilingual Terminology Extraction from Comparable Corpora
Comparable corpora are the main alternative to the use of parallel corpora to extract bilingual lexicons. Although it is easier to build comparable corpora, specialized comparable corpora are often of modest size in comp…
Machine TranslationTopic ModelsTranslationWord Embeddings