paper-with-me

홈 › Papers

Wikipedia Entities as Rendezvous across Languages: Grounding Multilingual Language Models by Predicting Wikipedia Hyperlinks

2021-06-01 · NAACL 2021 4 · Iacer Calixto, Alessandro Raganato, Tommaso Pasini

Masked language models have quickly become the de facto standard when processing text. Recently, several approaches have been proposed to further enrich word representations with external knowledge sources such as knowledge graphs. However, these models are devised and evaluated in a monolingual setting only. In this work, we propose a language-independent entity prediction task as an intermediate training procedure to ground word representations on entity semantics and bridge the gap across different languages by means of a shared vocabulary of entities. We show that our approach effectively injects new lexical-semantic knowledge into neural models, improving their performance on different semantic tasks in the zero-shot crosslingual setting. As an additional advantage, our intermediate training does not require any supplementary input, allowing our models to be applied to new datasets right away. In our experiments, we use Wikipedia articles in up to 100 languages and already observe consistent gains compared to strong baselines when predicting entities using only the English Wikipedia. Further adding extra languages lead to improvements in most tasks up to a certain point, but overall we found it non-trivial to scale improvements in model transferability by training on ever increasing amounts of Wikipedia languages.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesKnowledge Graphs

Similar Papers 제목 키워드 기반

Illinois Cross-Lingual Wikifier: Grounding Entities in Many Languages to the English Wikipedia

2016-12-01 · COLING 2016 12 · Chen-Tse Tsai, Dan Roth

We release a cross-lingual wikification system for all languages in Wikipedia. Given a piece of text in any supported language, the system identifies names of people, locations, organizations, and grounds these names to …

Cross-Lingual NEREntity Linkingnamed-entity-recognitionNamed Entity Recognition+2

Design Challenges in Low-resource Cross-lingual Entity Linking

2020-05-02 · EMNLP 2020 11 · Xingyu Fu, Weijia Shi, Xiaodong Yu, Zian Zhao 외

Cross-lingual Entity Linking (XEL), the problem of grounding mentions of entities in a foreign language text into an English knowledge base such as Wikipedia, has seen a lot of research in recent years, with a range of p…

Cross-Lingual Entity LinkingEntity Linking

Pairing Wikipedia Articles Across Languages

2016-12-01 · WS 2016 12 · Marcus Klang, Pierre Nugues

Wikipedia has become a reference knowledge source for scores of NLP applications. One of its invaluable features lies in its multilingual nature, where articles on a same entity or concept can have from one to more than …

ArticlesQuestion AnsweringWord Alignment

Global Entity Ranking Across Multiple Languages

2017-03-17 · Prantik Bhattacharyya, Nemanja Spasojevic

We present work on building a global long-tailed ranking of entities across multiple languages using Wikipedia and Freebase knowledge bases. We identify multiple features and build a model to rank entities using a ground…

A Parallel English - Serbian - Bulgarian - Macedonian Lexicon of Named Entities

2022-09-01 · CLIB 2022 9 · Aleksandar Petrovski

This paper describes the creation of a parallel multilingual lexicon of named entities from English to three South Slavic languages: Serbian, Bulgarian and Macedonian, with Wikipedia as a source. The basics of the propos…

Miscellaneous