paper-with-me

홈 › Papers

The Dutch LESLLA Corpus

2014-05-01 · LREC 2014 5 · S, Eric ers, Ineke van de Craats, Vanja de Lint

This paper describes the Dutch LESLLA data and its curation. LESLLA stands for Low-Educated Second Language and Literacy Acquisition. The data was collected for research in this field and would have been disappeared if it were not saved. Within the CLARIN project Data Curation Service the data was made into a spoken language resource and made available to other researchers.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language Acquisition

Similar Papers 제목 키워드 기반

GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training

2026-04-01 · Jesse van Oort, Frank Brinkkemper, Erik de Graaf, Bram Vanroy 외 arxiv

We present the GPT-NL Public Corpus, the biggest permissively licensed corpus of Dutch language resources. The GPT-NL Public Corpus contains 21 Dutch-only collections totalling 36B preprocessed Dutch tokens not present i…

DutchSemCor: Targeting the ideal sense-tagged corpus

2012-05-01 · LREC 2012 5 · Piek Vossen, Attila G{\"o}r{\"o}g, Rub{\'e}n Izquierdo, Antal Van den Bosch

Word Sense Disambiguation (WSD) systems require large sense-tagged corpora along with lexical databases to reach satisfactory results. The number of English language resources for developed WSD increased in the past year…

Active LearningWord Sense Disambiguation

Language corpora for the Dutch medical domain

2026-04-28 · B. van Es arxiv

\textbf{Background:} Dutch medical corpora are scarce, limiting NLP development. \\ \textbf{Methods:} We translated English datasets, identified medical text in generic corpora, and extracted open Dutch medical resources…

Negation Detection in Dutch Spoken Human-Computer Conversations

2022-06-01 · LREC 2022 6 · Tom Sweers, Iris Hendrickx, Helmer Strik

Proper recognition and interpretation of negation signals in text or communication is crucial for any form of full natural language understanding. It is also essential for computational approaches to natural language pro…

Natural Language UnderstandingNegationNegation DetectionTransfer Learning

Biomedical Entity Linking for Dutch: Fine-tuning a Self-alignment BERT Model on an Automatically Generated Wikipedia Corpus

2024-05-20 · Fons Hartendorp, Tom Seinen, Erik van Mulligen, Suzan Verberne

Biomedical entity linking, a main component in automatic information extraction from health-related texts, plays a pivotal role in connecting textual entities (such as diseases, drugs and body parts mentioned by patients…

Entity Linking