paper-with-me

Papers

Multilingual Abusiveness Identification on Code-Mixed Social Media Text

2022-03-01 · Ekagra Ranjan, Naman Poddar

Social Media platforms have been seeing adoption and growth in their usage over time. This growth has been further accelerated with the lockdown in the past year when people's interaction, conversation, and expression were limited physically. It is becoming increasingly important to keep the platform safe from abusive content for better user experience. Much work has been done on English social media content but text analysis on non-English social media is relatively underexplored. Non-English social media content have the additional challenges of code-mixing, transliteration and using different scripture in same sentence. In this work, we propose an approach for abusiveness identification on the multilingual Moj dataset which comprises of Indic languages. Our approach tackles the common challenges of non-English social media content and can be extended to other languages as well.

📄 PDF Abstract BibTeX arXiv:2204.01848

Code (0)

등록된 구현이 없습니다.

Tasks

SentenceTransliteration

Similar Papers 제목 키워드 기반

A Fast, Compact, Accurate Model for Language Identification of Codemixed Text

2018-10-09 · EMNLP 2018 10 · Yuan Zhang, Jason Riesa, Daniel Gillick, Anton Bakalov 외

We address fine-grained multilingual language identification: providing a language code for every token in a sentence, including codemixed text containing multiple languages. Such text is prevalent online, in documents, …

DecoderLanguage IdentificationSentence

OFFLangOne@DravidianLangTech-EACL2021: Transformers with the Class Balanced Loss for Offensive Language Identification in Dravidian Code-Mixed text.

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Suman Dowlagar, Radhika Mamidi

The intensity of online abuse has increased in recent years. Automated tools are being developed to prevent the use of hate speech and offensive content. Most of the technologies use natural language and machine learning…

Language IdentificationTransliteration

MMT: A Multilingual and Multi-Topic Indian Social Media Dataset

2023-04-02 · Dwip Dalal, Vivek Srivastava, Mayank Singh

Social media plays a significant role in cross-cultural communication. A vast amount of this occurs in code-mixed and multilingual form, posing a significant challenge to Natural Language Processing (NLP) tools for proce…

DiversityLanguage Identificationnamed-entity-recognitionNamed Entity Recognition

JUNLP@DravidianLangTech-EACL2021: Offensive Language Identification in Dravidian Langauges

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Avishek Garain, Atanu Mandal, Sudip Kumar Naskar

Offensive language identification has been an active area of research in natural language processing. With the emergence of multiple social media platforms offensive language identification has emerged as a need of the h…

Language Identification

Towards Offensive Language Identification for Dravidian Languages

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Siva Sai, Yashvardhan Sharma

Offensive speech identification in countries like India poses several challenges due to the usage of code-mixed and romanized variants of multiple languages by the users in their posts on social media. The challenge of o…

Few-Shot LearningLanguage IdentificationTransfer LearningTransliteration+1