Exploring Classifier Combinations for Language Variety Identification
This paper describes CLiPS{'}s submissions for the Discriminating between Dutch and Flemish in Subtitles (DFS) shared task at VarDial 2018. We explore different ways to combine classifiers trained on different feature groups. Our best system uses two Linear SVM classifiers; one trained on lexical features (word n-grams) and one trained on syntactic features (PoS n-grams). The final prediction for a document to be in Flemish Dutch or Netherlandic Dutch is made by the classifier that outputs the highest probability for one of the two labels. This confidence vote approach outperforms a meta-classifier on the development data and on the test data.
Code (0)
등록된 구현이 없습니다.
Tasks
Language IdentificationPOSMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Exploring Lexical and Syntactic Features for Language Variety Identification
We present a method to discriminate between texts written in either the Netherlandic or the Flemish variant of the Dutch language. The method draws on a feature bundle representing text statistics, syntactic features, an…
BIG-bench Machine LearningLanguage IdentificationSentenceText ClassificationIncluding Dialects and Language Varieties in Author Profiling
This paper presents a computational approach to author profiling taking gender and language variety into account. We apply an ensemble system with the output of multiple linear SVM classifiers trained on character and wo…
Author ProfilingClassifier Ensembles for Dialect and Language Variety Identification
In this paper we present ensemble-based systems for dialect and language variety identification using the datasets made available by the organizers of the VarDial Evaluation Campaign 2018. We present a system developed t…
Dialect IdentificationTeam Ned Leeds at SemEval-2019 Task 4: Exploring Language Indicators of Hyperpartisan Reporting
This paper reports an experiment carried out to investigate the relevance of several syntactic, stylistic and pragmatic features on the task of distinguishing between mainstream and partisan news articles. The results of…
ArticlesAutomatic Token and Turn Level Language Identification for Code-Switched Text Dialog: An Analysis Across Language Pairs and Corpora
We examine the efficacy of various feature{--}learner combinations for language identification in different types of text-based code-switched interactions {--} human-human dialog, human-machine dialog as well as monolog …
Language IdentificationSpoken Language UnderstandingText Generation