Towards Offensive Language Identification for Tamil Code-Mixed YouTube Comments and Posts
Offensive Language detection in social media platforms has been an active field of research over the past years. In non-native English spoken countries, social media users mostly use a code-mixed form of text in their posts/comments. This poses several challenges in the offensive content identification tasks, and considering the low resources available for Tamil, the task becomes much harder. The current study presents extensive experiments using multiple deep learning, and transfer learning models to detect offensive content on YouTube. We propose a novel and flexible approach of selective translation and transliteration techniques to reap better results from fine-tuning and ensembling multilingual transformer networks like BERT, Distil- BERT, and XLM-RoBERTa. The experimental results showed that ULMFiT is the best model for this task. The best performing models were ULMFiT and mBERTBiLSTM for this Tamil code-mix dataset instead of more popular transfer learning models such as Distil- BERT and XLM-RoBERTa and hybrid deep learning models. The proposed model ULMFiT and mBERTBiLSTM yielded good results and are promising for effective offensive speech identification in low-resourced languages.
Code (1)
Tasks
Language IdentificationTransfer LearningTranslationTransliterationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
DOSA: Dravidian Code-Mixed Offensive Span Identification Dataset
This paper presents the Dravidian Offensive Span Identification Dataset (DOSA) for under-resourced Tamil-English and Kannada-English code-mixed text. The dataset addresses the lack of code-mixed datasets with annotated o…
Language IdentificationOffTamil@DravideanLangTech-EASL2021: Offensive Language Identification in Tamil Text
In the last few decades, Code-Mixed Offensive texts are used penetratingly in social media posts. Social media platforms and online communities showed much interest on offensive text identification in recent years. Conse…
feature selectionLanguage IdentificationFindings of the Shared Task on Offensive Span Identification from Code-Mixed Tamil-English Comments
Offensive content moderation is vital in social media platforms to support healthy online discussions. However, their prevalence in codemixed Dravidian languages is limited to classifying whole comments without identifyi…
Findings of the Shared Task on Offensive Span Identification fromCode-Mixed Tamil-English Comments
Offensive content moderation is vital in social media platforms to support healthy online discussions. However, their prevalence in code-mixed Dravidian languages is limited to classifying whole comments without identify…
JUNLP@DravidianLangTech-EACL2021: Offensive Language Identification in Dravidian Langauges
Offensive language identification has been an active area of research in natural language processing. With the emergence of multiple social media platforms offensive language identification has emerged as a need of the h…
Language Identification