paper-with-me

홈 › Papers

Exploring Classifier Combinations for Language Variety Identification

2018-08-01 · COLING 2018 8 · Tim Kreutz, Walter Daelemans

This paper describes CLiPS{'}s submissions for the Discriminating between Dutch and Flemish in Subtitles (DFS) shared task at VarDial 2018. We explore different ways to combine classifiers trained on different feature groups. Our best system uses two Linear SVM classifiers; one trained on lexical features (word n-grams) and one trained on syntactic features (PoS n-grams). The final prediction for a document to be in Flemish Dutch or Netherlandic Dutch is made by the classifier that outputs the highest probability for one of the two labels. This confidence vote approach outperforms a meta-classifier on the development data and on the test data.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Language IdentificationPOS

Methods 이 논문이 사용한 방법론

SVM A Support Vector Machine, or SVM, is a non-parametric supervised learning model. For non-linear classification and regression, they utilise the kernel trick to map inputs…

Similar Papers 제목 키워드 기반

Exploring Lexical and Syntactic Features for Language Variety Identification

2017-04-01 · WS 2017 4 · Chris van der Lee, Antal Van den Bosch

We present a method to discriminate between texts written in either the Netherlandic or the Flemish variant of the Dutch language. The method draws on a feature bundle representing text statistics, syntactic features, an…

BIG-bench Machine LearningLanguage IdentificationSentenceText Classification

Including Dialects and Language Varieties in Author Profiling

2017-07-03 · Alina Maria Ciobanu, Marcos Zampieri, Shervin Malmasi, Liviu P. Dinu

This paper presents a computational approach to author profiling taking gender and language variety into account. We apply an ensemble system with the output of multiple linear SVM classifiers trained on character and wo…

Author Profiling

Classifier Ensembles for Dialect and Language Variety Identification

2018-08-14 · Liviu P. Dinu, Alina Maria Ciobanu, Marcos Zampieri, Shervin Malmasi

In this paper we present ensemble-based systems for dialect and language variety identification using the datasets made available by the organizers of the VarDial Evaluation Campaign 2018. We present a system developed t…

Dialect Identification

Team Ned Leeds at SemEval-2019 Task 4: Exploring Language Indicators of Hyperpartisan Reporting

2019-06-01 · SEMEVAL 2019 6 · Bozhidar Stevanoski, Sonja Gievska

This paper reports an experiment carried out to investigate the relevance of several syntactic, stylistic and pragmatic features on the task of distinguishing between mainstream and partisan news articles. The results of…

Articles

Automatic Token and Turn Level Language Identification for Code-Switched Text Dialog: An Analysis Across Language Pairs and Corpora

2018-07-01 · WS 2018 7 · Vikram Ramanarayanan, Robert Pugh

We examine the efficacy of various feature{--}learner combinations for language identification in different types of text-based code-switched interactions {--} human-human dialog, human-machine dialog as well as monolog …

Language IdentificationSpoken Language UnderstandingText Generation