paper-with-me

홈 › Papers

CLiPS Stylometry Investigation (CSI) corpus: A Dutch corpus for the detection of age, gender, personality, sentiment and deception in text

2014-05-01 · LREC 2014 5 · Ben Verhoeven, Walter Daelemans

We present the CLiPS Stylometry Investigation (CSI) corpus, a new Dutch corpus containing reviews and essays written by university students. It is designed to serve multiple purposes: detection of age, gender, authorship, personality, sentiment, deception, topic and genre. Another major advantage is its planned yearly expansion with each year{'}s new students. The corpus currently contains about 305,000 tokens spread over 749 documents. The average review length is 128 tokens; the average essay length is 1126 tokens. The corpus will be made available on the CLiPS website (www.clips.uantwerpen.be/datasets) and can freely be used for academic research purposes. An initial deception detection experiment was performed on this data. Deception detection is the task of automatically classifying a text as being either truthful or deceptive, in our case by examining the writing style of the author. This task has never been investigated for Dutch before. We performed a supervised machine learning experiment using the SVM algorithm in a 10-fold cross-validation setup. The only features were the token unigrams present in the training data. Using this simple method, we reached a state-of-the-art F-score of 72.2{\%}.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Deception DetectionText Classification

Similar Papers 제목 키워드 기반

TwiSty: A Multilingual Twitter Stylometry Corpus for Gender and Personality Profiling

2016-05-01 · LREC 2016 5 · Ben Verhoeven, Walter Daelemans, Barbara Plank

Personality profiling is the task of detecting personality traits of authors based on writing style. Several personality typologies exist, however, the Briggs-Myer Type Indicator (MBTI) is particularly popular in the non…

Gender Prediction

GPT-NL Public Corpus: A Permissively Licensed, Dutch-First Dataset for LLM Pre-training

2026-04-01 · Jesse van Oort, Frank Brinkkemper, Erik de Graaf, Bram Vanroy 외 arxiv

We present the GPT-NL Public Corpus, the biggest permissively licensed corpus of Dutch language resources. The GPT-NL Public Corpus contains 21 Dutch-only collections totalling 36B preprocessed Dutch tokens not present i…

A Dictionary-based Approach to Racism Detection in Dutch Social Media

2016-08-31 · Stéphan Tulkens, Lisa Hilte, Elise Lodewyckx, Ben Verhoeven 외

We present a dictionary-based approach to racism detection in Dutch social media comments, which were retrieved from two public Belgian social media sites likely to attract racist reactions. These comments were labeled a…

DutchSemCor: Targeting the ideal sense-tagged corpus

2012-05-01 · LREC 2012 5 · Piek Vossen, Attila G{\"o}r{\"o}g, Rub{\'e}n Izquierdo, Antal Van den Bosch

Word Sense Disambiguation (WSD) systems require large sense-tagged corpora along with lexical databases to reach satisfactory results. The number of English language resources for developed WSD increased in the past year…

Active LearningWord Sense Disambiguation

Language corpora for the Dutch medical domain

2026-04-28 · B. van Es arxiv

\textbf{Background:} Dutch medical corpora are scarce, limiting NLP development. \\ \textbf{Methods:} We translated English datasets, identified medical text in generic corpora, and extracted open Dutch medical resources…