Curation of Dutch Regional Dictionaries
This paper describes the process of semi-automatically converting dictionaries from paper to structured text (database) and the integration of these into the CLARIN infrastructure in order to make the dictionaries accessible and retrievable for the research community. The case study at hand is that of the curation of 42 fascicles of the Dictionaries of the Brabantic and Limburgian dialects, and 6 fascicles of the Dictionary of dialects in Gelderland.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
The Dutch LESLLA Corpus
This paper describes the Dutch LESLLA data and its curation. LESLLA stands for Low-Educated Second Language and Literacy Acquisition. The data was collected for research in this field and would have been disappeared if i…
Language AcquisitionA Dictionary-based Approach to Racism Detection in Dutch Social Media
We present a dictionary-based approach to racism detection in Dutch social media comments, which were retrieved from two public Belgian social media sites likely to attract racist reactions. These comments were labeled a…
Quantifying Language Variation Acoustically with Few Resources
Deep acoustic models represent linguistic information based on massive amounts of data. Unfortunately, for regional languages and dialects such resources are mostly not available. However, deep acoustic models might have…
Dynamic Time WarpingAn Extensible Multilingual Open Source Lemmatizer
We present GATE DictLemmatizer, a multilingual open source lemmatizer for the GATE NLP framework that currently supports English, German, Italian, French, Dutch, and Spanish, and is easily extensible to other languages. …
Information RetrievalLEMMALemmatizationA Longitudinal Bilingual Frisian-Dutch Radio Broadcast Database Designed for Code-Switching Research
We present a new speech database containing 18.5 hours of annotated radio broadcasts in the Frisian language. Frisian is mostly spoken in the province Fryslan and it is the second official language of the Netherlands. Th…