Overview of the Fourth BUCC Shared Task: Bilingual Dictionary Induction from Comparable Corpora
The shared task of the 13th Workshop on Building and Using Comparable Corpora was devoted to the induction of bilingual dictionaries from comparable rather than parallel corpora. In this task, for a number of language pairs involving Chinese, English, French, German, Russian and Spanish, the participants were supposed to determine automatically the target language translations of several thousand source language test words of three frequency ranges. We describe here some background, the task definition, the training and test data sets and the evaluation used for ranking the participating systems. We also summarize the approaches used and present the results of the evaluation. In conclusion, the outcome of the competition are the results of a number of systems which provide surprisingly good solutions to the ambitious problem.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
BUCC2020: Bilingual Dictionary Induction using Cross-lingual Embedding
This paper presents a deep learning system for the BUCC 2020 shared task: Bilingual dictionary induction from comparable corpora. We have submitted two runs for this shared Task, German (de) and English (en) language pai…
Deep LearningWord EmbeddingsCUNI Submission to the BUCC 2022 Shared Task on Bilingual Term Alignment
We present our submission to the BUCC Shared Task on bilingual term alignment in comparable specialized corpora. We devised three approaches using static embeddings with post-hoc alignment, the Monoses pipeline for unsup…
Machine TranslationTranslationLMU Bilingual Dictionary Induction System with Word Surface Similarity Scores for BUCC 2020
The task of Bilingual Dictionary Induction (BDI) consists of generating translations for source language words which is important in the framework of machine translation (MT). The aim of the BUCC 2020 shared task is to p…
Machine TranslationTranslationWord EmbeddingsWord SimilarityTALN/LS2N Participation at the BUCC Shared Task: Bilingual Dictionary Induction from Comparable Corpora
This paper describes the TALN/LS2N system participation at the Building and Using Comparable Corpora (BUCC) shared task. We first introduce three strategies: (i) a word embedding approach based on fastText embeddings; (i…
Overview of the Second BUCC Shared Task: Spotting Parallel Sentences in Comparable Corpora
This paper presents the BUCC 2017 shared task on parallel sentence extraction from comparable corpora. It recalls the design of the datasets, presents their final construction and statistics and the methods used to evalu…
Machine TranslationSentence