paper-with-me

홈 › Papers

This Table is Different: A WordNet-Based Approach to Identifying References to Document Entities

2016-01-01 · GWC 2016 1 · Shomir Wilson, Alan Black, Jon Oberlander

Writing intended to inform frequently contains references to document entities (DEs), a mixed class that includes orthographically structured items (e.g., illustrations, sections, lists) and discourse entities (arguments, suggestions, points). Such references are vital to the interpretation of documents, but they often eschew identifiers such as “Figure 1” for inexplicit phrases like “in this figure” or “from these premises”. We examine inexplicit references to DEs, termed DE references, and recast the problem of their automatic detection into the determination of relevant word senses. We then show the feasibility of machine learning for the detection of DE-relevant word senses, using a corpus of human-labeled synsets from WordNet. We test cross-domain performance by gathering lemmas and synsets from three corpora: website privacy policies, Wikipedia articles, and Wikibooks textbooks. Identifying DE references will enable language technologies to use the information encoded by them, permitting the automatic generation of finely-tuned descriptions of DEs and the presentation of richly-structured information to readers.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Articles

Similar Papers 제목 키워드 기반

Exploiting WordNet Synset and Hypernym Representations for Answer Selection

2020-12-01 · Asian Chapter of the Association for Computational Linguistics 2020 · Weikang Li, Yunfang Wu

Answer selection (AS) is an important subtask of document-based question answering (DQA). In this task, the candidate answers come from the same document, and each answer sentence is semantically related to the given que…

Answer SelectionQuestion AnsweringSentence

Text Document Clustering: Wordnet vs. TF-IDF vs. Word Embeddings

2021-01-01 · EACL (GWC) 2021 1 · Michał Marcińczuk, Mateusz Gniewkowski, Tomasz Walkowiak, Marcin Będkowski

In the paper, we deal with the problem of unsupervised text document clustering for the Polish language. Our goal is to compare the modern approaches based on language modeling (doc2vec and BERT) with the classical ones,…

ClusteringLanguage ModelingLanguage ModellingWord Embeddings

Using Arabic Wordnet for semantic indexation in information retrieval system

2013-06-11 · Mohammed Alaeddine Abderrahim, Mohammed El Amine Abderrahim, Mohammed Amine Chikh

In the context of arabic Information Retrieval Systems (IRS) guided by arabic ontology and to enable those systems to better respond to user requirements, this paper aims to representing documents and queries by the best…

Information RetrievalRetrieval

Adding Pronunciation Information to Wordnets

2020-05-01 · LREC 2020 5 · Thierry Declerck, Lenka Bajcetic, Melanie Siegel

We describe on-going work consisting in adding pronunciation information to wordnets, as such information can indicate specific senses of a word. Many wordnets associate with their senses only a lemma form and a part-of-…

LEMMATAG

Extending and Improving Wordnet via Unsupervised Word Embeddings

2017-04-29 · Mikhail Khodak, Andrej Risteski, Christiane Fellbaum, Sanjeev Arora

This work presents an unsupervised approach for improving WordNet that builds upon recent advances in document and sense representation via distributional semantics. We apply our methods to construct Wordnets in French a…

ClusteringWord Embeddings