Cross-Lingual Link Discovery for Under-Resourced Languages
In this paper, we provide an overview of current technologies for cross-lingual link discovery, and we discuss challenges, experiences and prospects of their application to under-resourced languages. We rst introduce the goals of cross-lingual linking and associated technologies, and in particular, the role that the Linked Data paradigm (Bizer et al., 2011) applied to language data can play in this context. We de ne under-resourced languages with a speci c focus on languages actively used on the internet, i.e., languages with a digitally versatile speaker community, but limited support in terms of language technology. We argue that languages for which considerable amounts of textual data and (at least) a bilingual word list are available, techniques for cross-lingual linking can be readily applied, and that these enable the implementation of downstream applications for under-resourced languages via the localisation and adaptation of existing technologies and resources.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Mahānāma: A Unique Testbed for Literary Entity Discovery and Linking
High lexical variation, ambiguous references, and long-range dependencies make entity resolution in literary texts particularly challenging. We present Mahānāma, the first large-scale dataset for end-to-end Entity Discov…
Entity ResolutionEntity LinkingmRAKL: Multilingual Retrieval-Augmented Knowledge Graph Construction for Low-Resourced Languages
Knowledge Graphs represent real-world entities and the relationships between them. Multilingual Knowledge Graph Construction (mKGC) refers to the task of automatically constructing or predicting missing entities and link…
Cross-Lingual TransferQuestion AnsweringKnowledge GraphsBeyond Weaponization: NLP Security for Medium and Lower-Resourced Languages in Their Own Right
Despite mounting evidence that multilinguality can be easily weaponized against language models (LMs), works across NLP Security remain overwhelmingly English-centric. In terms of securing LMs, the NLP norm of "English f…
How Does Language Influence Documentation Workflow? Unsupervised Word Discovery Using Translations in Multiple Languages
For language documentation initiatives, transcription is an expensive resource: one minute of audio is estimated to take one hour and a half on average of a linguist's work (Austin and Sallabank, 2013). Recently, collect…
Investigating the Quality of Static Anchor Embeddings from Transformers for Under-Resourced Languages
This paper reports on experiments for cross-lingual transfer using the anchor-based approach of Schuster et al. (2019) for English and a low-resourced language, namely Hindi. For the sake of comparison, we also evaluate …
Cross-Lingual Transfer