paper-with-me

Papers

Unseen Word Representation by Aligning Heterogeneous Lexical Semantic Spaces

2018-11-12 · Victor Prokhorov, Mohammad Taher Pilehvar, Dimitri Kartsaklis, Pietro Lio, Nigel Collier

Word embedding techniques heavily rely on the abundance of training data for individual words. Given the Zipfian distribution of words in natural language texts, a large number of words do not usually appear frequently or at all in the training data. In this paper we put forward a technique that exploits the knowledge encoded in lexical resources, such as WordNet, to induce embeddings for unseen words. Our approach adapts graph embedding and cross-lingual vector space transformation techniques in order to merge lexical knowledge encoded in ontologies with that derived from corpus statistics. We show that the approach can provide consistent performance improvements across multiple evaluation benchmarks: in-vitro, on multiple rare word similarity datasets, and in-vivo, in two downstream text classification tasks.

📄 PDF Abstract BibTeX arXiv:1811.04983

Code (0)

등록된 구현이 없습니다.

Tasks

General ClassificationGraph Embeddingtext-classificationText ClassificationWord Similarity

Similar Papers 제목 키워드 기반

Post-Specialisation: Retrofitting Vectors of Words Unseen in Lexical Resources

2018-05-08 · NAACL 2018 6 · Ivan Vulić, Goran Glavaš, Nikola Mrkšić, Anna Korhonen

Word vector specialisation (also known as retrofitting) is a portable, light-weight approach to fine-tuning arbitrary distributional word vector spaces by injecting external knowledge from rich lexical resources such as …

Dialogue State TrackingText SimplificationWord Similarity

A Robust Approach to Aligning Heterogeneous Lexical Resources

2014-06-01 · ACL 2014 6 · Mohammad Taher Pilehvar, Roberto Navigli
Semantic ParsingSemantic Role LabelingSemantic Textual SimilarityWord Sense Disambiguation

TUDA-CCL at SemEval-2021 Task 1: Using Gradient-boosted Regression Tree Ensembles Trained on a Heterogeneous Feature Set for Predicting Lexical Complexity

2021-08-01 · SEMEVAL 2021 · Sebastian Gombert, Sabine Bartsch

In this paper, we present our systems submitted to SemEval-2021 Task 1 on lexical complexity prediction.The aim of this shared task was to create systems able to predict the lexical complexity of word tokens and bigram m…

SentenceWord Embeddings

Locally Measuring Cross-lingual Lexical Alignment: A Domain and Word Level Perspective

2024-10-07 · Taelin Karidi, Eitan Grossman, Omri Abend

NLP research on aligning lexical representation spaces to one another has so far focused on aligning language spaces in their entirety. However, cognitive science has long focused on a local perspective, investigating wh…

Inducing Embeddings for Rare and Unseen Words by Leveraging Lexical Resources

2017-04-01 · EACL 2017 4 · Mohammad Taher Pilehvar, Nigel Collier

We put forward an approach that exploits the knowledge encoded in lexical resources in order to induce representations for words that were not encountered frequently during training. Our approach provides an advantage ov…

Word Embeddings