paper-with-me

홈 › Papers

The Unreasonable Effectiveness of Machine Learning in Moldavian versus Romanian Dialect Identification

2020-07-30 · Mihaela Găman, Radu Tudor Ionescu

Motivated by the seemingly high accuracy levels of machine learning models in Moldavian versus Romanian dialect identification and the increasing research interest on this topic, we provide a follow-up on the Moldavian versus Romanian Cross-Dialect Topic Identification (MRC) shared task of the VarDial 2019 Evaluation Campaign. The shared task included two sub-task types: one that consisted in discriminating between the Moldavian and Romanian dialects and one that consisted in classifying documents by topic across the two dialects of Romanian. Participants achieved impressive scores, e.g. the top model for Moldavian versus Romanian dialect identification obtained a macro F1 score of 0.895. We conduct a subjective evaluation by human annotators, showing that humans attain much lower accuracy rates compared to machine learning (ML) models. Hence, it remains unclear why the methods proposed by participants attain such high accuracy rates. Our goal is to understand (i) why the proposed methods work so well (by visualizing the discriminative features) and (ii) to what extent these methods can keep their high accuracy levels, e.g. when we shorten the text samples to single sentences or when we use tweets at inference time. A secondary goal of our work is to propose an improved ML model using ensemble learning. Our experiments show that ML models can accurately identify the dialects, even at the sentence level and across different domains (news articles versus tweets). We also analyze the most discriminative features of the best performing models, providing some explanations behind the decisions taken by these models. Interestingly, we learn new dialectal patterns previously unknown to us or to our human annotators. Furthermore, we conduct experiments showing that the machine learning performance on the MRC shared task can be improved through an ensemble based on stacking.

📄 PDF Abstract BibTeX arXiv:2007.15700

Code (1)

raduionescu/MOROCO-Tweets 공식 구현

Tasks

ArticlesBIG-bench Machine LearningDialect IdentificationEnsemble LearningSentence

Similar Papers 제목 키워드 기반

MOROCO: The Moldavian and Romanian Dialectal Corpus

2019-01-19 · ACL 2019 7 · Andrei M. Butnaru, Radu Tudor Ionescu

In this work, we introduce the MOldavian and ROmanian Dialectal COrpus (MOROCO), which is freely available for download at https://github.com/butnaruandrei/MOROCO. The corpus contains 33564 samples of text (with over 10 …

Cultural Vocal Bursts Intensity Prediction

DTeam @ VarDial 2019: Ensemble based on skip-gram and triplet loss neural networks for Moldavian vs. Romanian cross-dialect topic identification

2019-06-01 · WS 2019 6 · Diana Tudoreanu

This paper presents the solution proposed by DTeam in the VarDial 2019 Evaluation Campaign for the Moldavian vs. Romanian cross-topic identification task. The solution proposed is a Support Vector Machines (SVM) ensemble…

General ClassificationTriplet

Dialect Identification under Domain Shift: Experiments with Discriminating Romanian and Moldavian

2020-12-01 · VarDial (COLING) 2020 12 · Çağrı Çöltekin

This paper describes a set of experiments for discriminating between two closely related language varieties, Moldavian and Romanian, under a substantial domain shift. The experiments were conducted as part of the Romania…

Dialect IdentificationPosition

SC-UPB at the VarDial 2019 Evaluation Campaign: Moldavian vs. Romanian Cross-Dialect Topic Identification

2019-06-01 · WS 2019 6 · Cristian Onose, Dumitru-Clementin Cercel, Stefan Trausan-Matu

This paper describes our models for the Moldavian vs. Romanian Cross-Topic Identification (MRC) evaluation campaign, part of the VarDial 2019 workshop. We focus on the three subtasks for MRC: binary classification betwee…

Binary ClassificationGeneral ClassificationMulti-class Classification

Discriminating between standard Romanian and Moldavian tweets using filtered character ngrams

2020-12-01 · VarDial (COLING) 2020 12 · Andrea Ceolin, Hong Zhang

We applied word unigram models, character ngram models, and CNNs to the task of distinguishing tweets of two related dialects of Romanian (standard Romanian and Moldavian) for the VarDial 2020 RDI shared task (Gaman et a…

Articlestext-classificationText Classification