paper-with-me

홈 › Papers

DALC: the Dutch Abusive Language Corpus

2021-08-01 · ACL (WOAH) 2021 8 · Tommaso Caselli, Arjan Schelhaas, Marieke Weultjes, Folkert Leistra, Hylke van der Veen, Gerben Timmerman, Malvina Nissim

As socially unacceptable language become pervasive in social media platforms, the need for automatic content moderation become more pressing. This contribution introduces the Dutch Abusive Language Corpus (DALC v1.0), a new dataset with tweets manually an- notated for abusive language. The resource ad- dress a gap in language resources for Dutch and adopts a multi-layer annotation scheme modeling the explicitness and the target of the abusive messages. Baselines experiments on all annotation layers have been conducted, achieving a macro F1 score of 0.748 for binary classification of the explicitness layer and .489 for target classification.

📄 PDF Abstract BibTeX

Code (1)

tommasoc80/dalc 공식 구현

Tasks

Abusive LanguageBinary ClassificationClassification

Similar Papers 제목 키워드 기반

“Zo Grof !”: A Comprehensive Corpus for Offensive and Abusive Language in Dutch

2022-07-01 · NAACL (WOAH) 2022 7 · Ward Ruitenbeek, Victor Zwart, Robin Van Der Noord, Zhenja Gnezdilov 외

This paper presents a comprehensive corpus for the study of socially unacceptable language in Dutch. The corpus extends and revise an existing resource with more data and introduces a new annotation dimension for offensi…

Abusive LanguageBinary Classification

Abusive content detection in transliterated Bengali-English social media corpus

2021-06-01 · NAACL (CALCS) 2021 6 · Salim Sazzed

Abusive text detection in low-resource languages such as Bengali is a challenging task due to the inadequacy of resources and tools. The ubiquity of transliterated Bengali comments in social media makes the task even mor…

Text Detection

GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training

2026-04-01 · Jesse van Oort, Frank Brinkkemper, Erik de Graaf, Bram Vanroy 외 arxiv

We present the GPT-NL Public Corpus, the biggest permissively licensed corpus of Dutch language resources. The GPT-NL Public Corpus contains 21 Dutch-only collections totalling 36B preprocessed Dutch tokens not present i…

DutchSemCor: Targeting the ideal sense-tagged corpus

2012-05-01 · LREC 2012 5 · Piek Vossen, Attila G{\"o}r{\"o}g, Rub{\'e}n Izquierdo, Antal Van den Bosch

Word Sense Disambiguation (WSD) systems require large sense-tagged corpora along with lexical databases to reach satisfactory results. The number of English language resources for developed WSD increased in the past year…

Active LearningWord Sense Disambiguation

Language corpora for the Dutch medical domain

2026-04-28 · B. van Es arxiv

\textbf{Background:} Dutch medical corpora are scarce, limiting NLP development. \\ \textbf{Methods:} We translated English datasets, identified medical text in generic corpora, and extracted open Dutch medical resources…