paper-with-me

Papers

Comparing Similarity Measures for Distributional Thesauri

2014-05-01 · LREC 2014 5 · Muntsa Padr{\'o}, Marco Idiart, Aline Villavicencio, Carlos Ramisch

Distributional thesauri have been applied for a variety of tasks involving semantic relatedness. In this paper, we investigate the impact of three parameters: similarity measures, frequency thresholds and association scores. We focus on the robustness and stability of the resulting thesauri, measuring inter-thesaurus agreement when testing different parameter values. The results obtained show that low-frequency thresholds affect thesaurus quality more than similarity measures, with more agreement found for increasing thresholds.These results indicate the sensitivity of distributional thesauri to frequency. Nonetheless, the observed differences do not transpose over extrinsic evaluation using TOEFL-like questions. While this may be specific to the task, we argue that a careful examination of the stability of distributional resources prior to application is needed.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Dimensionality Reduction

Similar Papers 제목 키워드 기반

B2SG: a TOEFL-like Task for Portuguese

2016-05-01 · LREC 2016 5 · Rodrigo Wilkens, Leonardo Zilio, Eduardo Ferreira, Aline Villavicencio

Resources such as WordNet are useful for NLP applications, but their manual construction consumes time and personnel, and frequently results in low coverage. One alternative is the automatic construction of large resourc…

Learning Thesaurus Relations from Distributional Features

2016-05-01 · LREC 2016 5 · Rosa Tsegaye Aga, Christian Wartena, Lucas Drumond, Lars Schmidt-Thieme

In distributional semantics words are represented by aggregated context features. The similarity of words can be computed by comparing their feature vectors. Thus, we can predict whether two words are synonymous or simil…

Document ClassificationRelation

Distributed Distributional Similarities of Google Books Over the Centuries

2014-05-01 · LREC 2014 5 · Martin Riedl, Richard Steuer, Chris Biemann

This paper introduces a distributional thesaurus and sense clusters computed on the complete Google Syntactic N-grams, which is extracted from Google Books, a very large corpus of digitized books published between 1520 a…

Graph Clustering

Extrinsic Evaluation of French Dependency Parsers on a Specialized Corpus: Comparison of Distributional Thesauri

2020-05-01 · LREC 2020 5 · Ludovic Tanguy, Pauline Brunet, Olivier Ferret

We present a study in which we compare 11 different French dependency parsers on a specialized corpus (consisting of research articles on NLP from the proceedings of the TALN conference). Due to the lack of a suitable go…

Articles

Building Static Embeddings from Contextual Ones: Is It Useful for Building Distributional Thesauri?

2022-06-01 · LREC 2022 6 · Olivier Ferret

While contextual language models are now dominant in the field of Natural Language Processing, the representations they build at the token level are not always suitable for all uses. In this article, we propose a new met…

Semantic SimilaritySemantic Textual SimilarityVocal Bursts Type Prediction