Span Extraction Aided Improved Code-mixed Sentiment Classification
Sentiment classification is a fundamental NLP task of detecting the sentiment polarity of a given text. In this paper we show how solving sentiment span extraction as an auxiliary task can help improve final sentiment classification performance in a low-resource code-mixed setup. To be precise, we don’t solve a simple multi-task learning objective, but rather design a unified transformer framework that exploits the bidirectional connection between the two tasks simultaneously. To facilitate research in this direction we release gold-standard human-annotated sentiment span extraction dataset for Tamil-english code-switched texts. Extensive experiments and strong baselines show that our proposed approach outperforms sentiment and span prediction by 1.27% and 2.78% respectively when compared to the best performing MTL baseline. We also establish the generalizability of our approach on the Twitter Sentiment Extraction dataset. We make our code and data publicly available on GitHub
Code (1)
Tasks
ClassificationMulti-Task LearningSentiment AnalysisSentiment ClassificationSimilar Papers 제목 키워드 기반
FII-UAIC at SemEval-2020 Task 9: Sentiment Analysis for Code-Mixed Social Media Text Using CNN
The {``}Sentiment Analysis for Code-Mixed Social Media Text{''} task at the SemEval 2020 competition focuses on sentiment analysis in code-mixed social media text , specifically, on the combination of English with Spanis…
Sentiment AnalysisZero-shot Code-Mixed Offensive Span Identification through Rationale Extraction
This paper investigates the effectiveness of sentence-level transformers for zero-shot offensive span identification on a code-mixed Tamil dataset. More specifically, we evaluate rationale extraction methods of Local Int…
Data AugmentationSentenceA Simple and Efficient Probabilistic Language model for Code-Mixed Text
The conventional natural language processing approaches are not accustomed to the social media text due to colloquial discourse and non-homogeneous characteristics. Significantly, the language identification in a multili…
Information RetrievalLanguage IdentificationLanguage ModelingLanguage Modelling+6DOSA: Dravidian Code-Mixed Offensive Span Identification Dataset
This paper presents the Dravidian Offensive Span Identification Dataset (DOSA) for under-resourced Tamil-English and Kannada-English code-mixed text. The dataset addresses the lack of code-mixed datasets with annotated o…
Language IdentificationHAPS-assisted Hybrid RF-FSO Multicast Communications: Error and Outage Analysis
In this work, we study the performance of multiple-hop mixed frequency (RF)/free-space optical (FSO) communication-based decode-and-forward protocol for multicast networks. So far, serving a large number of users is cons…