Including Dialects and Language Varieties in Author Profiling
This paper presents a computational approach to author profiling taking gender and language variety into account. We apply an ensemble system with the output of multiple linear SVM classifiers trained on character and word $n$-grams. We evaluate the system using the dataset provided by the organizers of the 2017 PAN lab on author profiling. Our approach achieved 75% average accuracy on gender identification on tweets written in four languages and 97% accuracy on language variety identification for Portuguese.
Code (0)
등록된 구현이 없습니다.
Tasks
Author ProfilingMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Vulgaris: Analysis of a Corpus for Middle-Age Varieties of Italian Language
Italian is a Romance language that has its roots in Vulgar Latin. The birth of the modern Italian started in Tuscany around the 14th century, and it is mainly attributed to the works of Dante Alighieri, Francesco Petrarc…
N-GrAM: New Groningen Author-profiling Model
We describe our participation in the PAN 2017 shared task on Author Profiling, identifying authors' gender and language variety for English, Spanish, Arabic and Portuguese. We describe both the final, submitted system, a…
Author ProfilingmodelPOSArap-Tweet: A Large Multi-Dialect Twitter Corpus for Gender, Age and Language Variety Identification
In this paper, we present Arap-Tweet, which is a large-scale and multi-dialectal corpus of Tweets from 11 regions and 16 countries in the Arab world representing the major Arabic dialectal varieties. To build this corpus…
Author ProfilingSentence-level dialects identification in the greater China region
Identifying the different varieties of the same language is more challenging than unrelated languages identification. In this paper, we propose an approach to discriminate language varieties or dialects of Mandarin Chine…
SentenceWord AlignmentDiscrimination between Similar Languages, Varieties and Dialects using CNN- and LSTM-based Deep Neural Networks
In this paper, we describe a system (CGLI) for discriminating similar languages, varieties and dialects using convolutional neural networks (CNNs) and long short-term memory (LSTM) neural networks. We have participated i…
Dialect IdentificationInformation RetrievalLanguage IdentificationMachine Translation+2