Word2Vec vs DBnary: Augmenting METEOR using Vector Representations or Lexical Resources?
This paper presents an approach combining lexico-semantic resources and distributed representations of words applied to the evaluation in machine translation (MT). This study is made through the enrichment of a well-known MT evaluation metric: METEOR. This metric enables an approximate match (synonymy or morphological similarity) between an automatic and a reference translation. Our experiments are made in the framework of the Metrics task of WMT 2014. We show that distributed representations are a good alternative to lexico-semantic resources for MT evaluation and they can even bring interesting additional information. The augmented versions of METEOR, using vector representations, are made available on our Github page.
Code (1)
Tasks
Machine TranslationTranslationSimilar Papers 제목 키워드 기반
Word2Vec vs DBnary ou comment (r\'e)concilier repr\'esentations distribu\'ees et r\'eseaux lexico-s\'emantiques ? Le cas de l'\'evaluation en traduction automatique (Word2Vec vs DBnary or how to bring back together vector representations and lexical resources ? A case study for machine translation evaluation)
Cet article pr{\'e}sente une approche associant r{\'e}seaux lexico-s{\'e}mantiques et repr{\'e}sentations distribu{\'e}es de mots appliqu{\'e}e {\`a} l{'}{\'e}valuation de la traduction automatique. Cette {\'e}tude est f…
Machine TranslationDbnary: Wiktionary as a LMF based Multilingual RDF network
Contributive resources, such as wikipedia, have proved to be valuable in Natural Language Processing or Multilingual Information Retrieval applications.This article focusses on Wiktionary, the dictionary part of the coll…
Information RetrievalRetrievalSeVeN: Augmenting Word Embeddings with Unsupervised Relation Vectors
We present SeVeN (Semantic Vector Networks), a hybrid resource that encodes relationships between words in the form of a graph. Different from traditional semantic networks, these relations are represented as vectors in …
RelationText CategorizationWord EmbeddingsWord SimilarityOrthogonality Constrained Multi-Head Attention For Keyword Spotting
Multi-head attention mechanism is capable of learning various representations from sequential data while paying attention to different subsequences, e.g., word-pieces or syllables in a spoken word. From the subsequences,…
Keyword SpottingAugmenting Chinese WordNet semantic relations with contextualized embeddings
Constructing semantic relations in WordNet has been a labour-intensive task, especially in a dynamic and fast-changing language environment. Combined with recent advancements of contextualized embeddings, this paper prop…