paper-with-me

Papers

Annotating the MASC Corpus with BabelNet

2014-05-01 · LREC 2014 5 · Andrea Moro, Roberto Navigli, Francesco Maria Tucci, Rebecca J. Passonneau

In this paper we tackle the problem of automatically annotating, with both word senses and named entities, the MASC 3.0 corpus, a large English corpus covering a wide range of genres of written and spoken text. We use BabelNet 2.0, a multilingual semantic network which integrates both lexicographic and encyclopedic knowledge, as our sense/entity inventory together with its semantic structure, to perform the aforementioned annotation task. Word sense annotated corpora have been around for more than twenty years, helping the development of Word Sense Disambiguation algorithms by providing both training and testing grounds. More recently Entity Linking has followed the same path, with the creation of huge resources containing annotated named entities. However, to date, there has been no resource that contains both kinds of annotation. In this paper we present an automatic approach for performing this annotation, together with its output on the MASC corpus. We use this corpus because its goal of integrating different types of annotations goes exactly in our same direction. Our overall aim is to stimulate research on the joint exploitation and disambiguation of word senses and named entities. Finally, we estimate the quality of our annotations using both manually-tagged named entities and word senses, obtaining an accuracy of roughly 70{\%} for both named entities and word sense annotations.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Entity LinkingReading ComprehensionRelation ExtractionWord Sense Disambiguation

Similar Papers 제목 키워드 기반

Sememe Prediction for BabelNet Synsets using Multilingual and Multimodal Information

2022-03-14 · Findings (ACL) 2022 5 · Fanchao Qi, Chuancheng Lv, Zhiyuan Liu, Xiaojun Meng 외

In linguistics, a sememe is defined as the minimum semantic unit of languages. Sememe knowledge bases (KBs), which are built by manually annotating words with sememes, have been successfully applied to various NLP tasks.…

TMASC: Transmasculine Attitude and Speech Corpus

2026-06-15 · Sidney Wong arxiv

We introduce the Transmasculine Attitudes and Speech Corpus (TMASC), a multimodal corpus of 196 transmasculine individuals, including questionnaire responses and 66 audio recordings. The questionnaire includes items expl…

The MASC Word Sense Corpus

2012-05-01 · LREC 2012 5 · Rebecca J. Passonneau, Collin F. Baker, Christiane Fellbaum, Nancy Ide

The MASC project has produced a multi-genre corpus with multiple layers of linguistic annotation, together with a sentence corpus containing WordNet 3.1 sense tags for 1000 occurrences of each of 100 words produced by mu…

Sentence

The Multilingual Affective Soccer Corpus (MASC): Compiling a biased parallel corpus on soccer reportage in English, German and Dutch

2016-09-01 · WS 2016 9 · Nadine Braun, Martijn Goudbeek, Emiel Krahmer
Text Generation

From Gentlemen to Frontiermen: Masculine Formations in English-Language Fiction (1771--1930)

2026-07-03 · Rong Wang arxiv

Masculinity in nineteenth-century fiction is not a single ideal but a field of competing scripts. Drawing on 150 British and American canonical novels from the txtLAB Novel450 corpus, published between 1771 and 1930, thi…

Coreference Resolution