Using Social Networks to Improve Language Variety Identification with Neural Networks
We propose a hierarchical neural network model for language variety identification that integrates information from a social network. Recently, language variety identification has enjoyed heightened popularity as an advanced task of language identification. The proposed model uses additional texts from a social network to improve language variety identification from two perspectives. First, they are used to introduce the effects of homophily. Secondly, they are used as expanded training data for shared layers of the proposed model. By introducing information from social networks, the model improved its accuracy by 1.67-5.56. Compared to state-of-the-art baselines, these improved performances are better in English and comparable in Spanish. Furthermore, we analyzed the cases of Portuguese and Arabic when the model showed weak performances, and found that the effect of homophily is likely to be weak due to sparsity and noises compared to languages with the strong performances.
Code (0)
등록된 구현이 없습니다.
Tasks
Language IdentificationSimilar Papers 제목 키워드 기반
Experiments in Language Variety Geolocation and Dialect Identification
In this paper we describe the systems we used when participating in the VarDial Evaluation Campaign organized as part of the 7th workshop on NLP for similar languages, varieties and dialects. The shared tasks we particip…
Dialect IdentificationFindings of the VarDial Evaluation Campaign 2021
This paper describes the results of the shared tasks organized as part of the VarDial Evaluation Campaign 2021. The campaign was part of the eighth workshop on Natural Language Processing (NLP) for Similar Languages, Var…
Dialect IdentificationLanguage IdentificationThe performance of multiple language models in identifying offensive language on social media
Text classification is an important topic in the field of natural language processing. It has been preliminarily applied in information retrieval, digital library, automatic abstracting, text filtering, word semantic dis…
Information RetrievalRetrievaltext-classificationText ClassificationTweetNLP: Cutting-Edge Natural Language Processing for Social Media
In this paper we present TweetNLP, an integrated platform for Natural Language Processing (NLP) in social media. TweetNLP supports a diverse set of NLP tasks, including generic focus areas such as sentiment analysis and …
Language IdentificationNamed Entity RecognitionNamed Entity Recognition (NER)Sentiment AnalysisAuthor Profiling at PAN: from Age and Gender Identification to Language Variety Identification (invited talk)
Author profiling is the study of how language is shared by people, a problem of growing importance in applications dealing with security, in order to understand who could be behind an anonymous threat message, and market…
Author ProfilingMarketingSentiment Analysis