paper-with-me

Papers

Improving Information Extraction from Wikipedia Texts using Basic English

2016-05-01 · LREC 2016 5 · Teresa Rodr{\'\i}guez-Ferreira, Adri{\'a}n Rabad{\'a}n, Raquel Herv{\'a}s, Alberto D{\'\i}az

The aim of this paper is to study the effect that the use of Basic English versus common English has on information extraction from online resources. The amount of online information available to the public grows exponentially, and is potentially an excellent resource for information extraction. The problem is that this information often comes in an unstructured format, such as plain text. In order to retrieve knowledge from this type of text, it must first be analysed to find the relevant details, and the nature of the language used can greatly impact the quality of the extracted information. In this paper, we compare triplets that represent definitions or properties of concepts obtained from three online collaborative resources (English Wikipedia, Simple English Wikipedia and Simple English Wiktionary) and study the differences in the results when Basic English is used instead of common English. The results show that resources written in Basic English produce less quantity of triplets, but with higher quality.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cross-lingual Biography Enrichment via Claim Extraction and Alignment

2026-08-24 · Yifei Song, Ziyang Chen, Emil Sayilov, Claire Gardent arxiv

English Wikipedia is often treated as the default encyclopedic source, yet non-English Wikipedia editions can contain richer locally grounded information for long-tail figures. We study cross-lingual biography enrichment…

WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from Wikipedia

2019-07-10 · EACL 2021 2 · Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong 외

We present an approach based on multilingual sentence embeddings to automatically extract parallel sentences from the content of Wikipedia articles in 85 languages, including several dialects or low-resource languages. W…

ArticlesSentenceSentence Embeddings

WikiCoref: An English Coreference-annotated Corpus of Wikipedia Articles

2016-05-01 · LREC 2016 5 · Abbas Ghaddar, Phillippe Langlais

This paper presents WikiCoref, an English corpus annotated for anaphoric relations, where all documents are from the English version of Wikipedia. Our annotation scheme follows the one of OntoNotes with a few disparities…

Articlescoreference-resolutionCoreference Resolution

A Parallel English - Serbian - Bulgarian - Macedonian Lexicon of Named Entities

2022-09-01 · CLIB 2022 9 · Aleksandar Petrovski

This paper describes the creation of a parallel multilingual lexicon of named entities from English to three South Slavic languages: Serbian, Bulgarian and Macedonian, with Wikipedia as a source. The basics of the propos…

Miscellaneous

DBpedia NIF: Open, Large-Scale and Multilingual Knowledge Extraction Corpus

2018-12-26 · Milan Dojchinovski, Julio Hernandez, Markus Ackermann, Amit Kirschenbaum 외

In the past decade, the DBpedia community has put significant amount of effort on developing technical infrastructure and methods for efficient extraction of structured information from Wikipedia. These efforts have been…

Articles