paper-with-me

홈 › Papers

Crawling Under-Resourced Languages - a Portal for Community-Contributed Corpus Collection

2022-06-01 · DCLRL (LREC) 2022 6 · Erik Körner, Felix Helfer, Christopher Schröder, Thomas Eckart, Dirk Goldhahn

The “Web as corpus” paradigm opens opportunities for enhancing the current state of language resources for endangered and under-resourced languages. However, standard crawling strategies tend to overlook available resources of these languages in favor of already well-documented ones. Since 2016, the “Crawling Under-Resourced Languages” portal (CURL) has been contributing to bridging the gap between established crawling techniques and knowledge about relevant Web resources that is only available in the specific language communities. The aim of the CURL portal is to enlarge the amount of available text material for under-resourced languages thereby developing available datasets further and to use them as a basis for statistical evaluation and enrichment of already available resources. The application is currently provided and further developed as part of the thematic cluster “Non-Latin scripts and Under-resourced languages” in the German national research consortium Text+. In this context, its focus lies on the extraction of text material and statistical information for the data domain “Lexical resources”.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages

2022-06-01 · EAMT 2022 6 · Marta Bañón, Miquel Esplà-Gomis, Mikel L. Forcada, Cristian García-Romero 외

We introduce the project “MaCoCu: Massive collection and curation of monolingual and bilingual data: focus on under-resourced languages”, funded by the Connecting Europe Facility, which is aimed at building monolingual a…

The Growing Gains and Pains of Iterative Web Corpora Crawling: Insights from South Slavic CLASSLA-web 2.0 Corpora

2026-01-16 · Taja Kuzman Pungeršek, Peter Rupnik, Vít Suchomel, Nikola Ljubešić arxiv

Crawling national top-level domains has proven to be highly effective for collecting texts in less-resourced languages. This approach has been recently used for South Slavic languages and resulted in the largest general …

Cross-Lingual Link Discovery for Under-Resourced Languages

2022-06-01 · LREC 2022 6 · Michael Rosner, Sina Ahmadi, Elena-Simona Apostol, Julia Bosque-Gil 외

In this paper, we provide an overview of current technologies for cross-lingual link discovery, and we discuss challenges, experiences and prospects of their application to under-resourced languages. We rst introduce the…

Survey on Thai NLP Language Resources and Tools

2022-06-01 · LREC 2022 6 · Ratchakrit Arreerard, Stephen Mander, Scott Piao

Over the past decades, Natural Language Processing (NLP) research has been expanding to cover more languages. Recently particularly, NLP community has paid increasing attention to under-resourced languages. However, ther…

Survey

Balanced End-to-End Monolingual pre-training for Low-Resourced Indic Languages Code-Switching Speech Recognition

2021-06-10 · Amir Hussein, Shammur Chowdhury, Najim Dehak, Ahmed Ali

The success in designing Code-Switching (CS) ASR often depends on the availability of the transcribed CS resources. Such dependency harms the development of ASR in low-resourced languages such as Bengali and Hindi. In th…

Language Modellingspeech-recognitionSpeech RecognitionTransfer Learning+1