paper-with-me

홈 › Papers

Sense-Annotated Corpora for Word Sense Disambiguation in Multiple Languages and Domains

2020-05-01 · LREC 2020 5 · Bianca Scarlini, Tommaso Pasini, Roberto Navigli

The knowledge acquisition bottleneck problem dramatically hampers the creation of sense-annotated data for Word Sense Disambiguation (WSD). Sense-annotated data are scarce for English and almost absent for other languages. This limits the range of action of deep-learning approaches, which today are at the base of any NLP task and are hungry for data. We mitigate this issue and encourage further research in multilingual WSD by releasing to the NLP community five large datasets annotated with word-senses in five different languages, namely, English, French, Italian, German and Spanish, and 5 distinct datasets in English, each for a different semantic domain. We show that supervised WSD models trained on our data attain higher performance than when trained on other automatically-created corpora. We release all our data containing more than 15 million annotated instances in 5 different languages at http://trainomatic.org/onesec.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Word Sense Disambiguation

Similar Papers 제목 키워드 기반

ChiSense-12: An English Sense-Annotated Child-Directed Speech Corpus

2022-06-01 · LREC 2022 6 · Francesco Cabiddu, Lewis Bott, Gary Jones, Chiara Gambi

Language acquisition research has benefitted from the use of annotated corpora of child-directed speech to examine key questions about how children learn and process language in real-world contexts. However, a lack of se…

Language AcquisitionWord Sense Disambiguation

An Iterative Approach for Unsupervised Most Frequent Sense Detection using WordNet and Word Embeddings

2018-01-01 · GWC 2018 1 · Kevin Patel, Pushpak Bhattacharyya

Given a word, what is the most frequent sense in which it occurs in a given corpus? Most Frequent Sense (MFS) is a strong baseline for unsupervised word sense disambiguation. If we have large amounts of sense-annotated c…

Word EmbeddingsWord Sense Disambiguation

Don't Neglect the Obvious: On the Role of Unambiguous Words in Word Sense Disambiguation

2020-04-29 · EMNLP 2020 11 · Daniel Loureiro, Jose Camacho-Collados

State-of-the-art methods for Word Sense Disambiguation (WSD) combine two different features: the power of pre-trained language models and a propagation method to extend the coverage of such models. This propagation is ne…

Word Sense Disambiguation

Context-Aware Semantic Similarity Measurement for Unsupervised Word Sense Disambiguation

2023-05-05 · Jorge Martinez-Gil

The issue of word sense ambiguity poses a significant challenge in natural language processing due to the scarcity of annotated data to feed machine learning models to face the challenge. Therefore, unsupervised word sen…

Semantic SimilaritySemantic Textual SimilarityWord Sense Disambiguation

Persian SemCor: A Bag of Word Sense Annotated Corpus for the Persian Language

2021-01-01 · EACL (GWC) 2021 1 · Hossein Rouhizadeh, Mehrnoush Shamsfard, Mahdi Dehghan, Masoud Rouhizadeh

Supervised approaches usually achieve the best performance in the Word Sense Disambiguation problem. However, the unavailability of large sense annotated corpora for many low-resource languages make these approaches inap…

Word Sense Disambiguation