paper-with-me

홈 › Papers

OffTamil@DravideanLangTech-EASL2021: Offensive Language Identification in Tamil Text

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Disne Sivalingam, Sajeetha Thavareesan

In the last few decades, Code-Mixed Offensive texts are used penetratingly in social media posts. Social media platforms and online communities showed much interest on offensive text identification in recent years. Consequently, research community is also interested in identifying such content and also contributed to the development of corpora. Many publicly available corpora are there for research on identifying offensive text written in English language but rare for low resourced languages like Tamil. The first code-mixed offensive text for Dravidian languages are developed by shared task organizers which is used for this study. This study focused on offensive language identification on code-mixed low-resourced Dravidian language Tamil using four classifiers (Support Vector Machine, random forest, k- Nearest Neighbour and Naive Bayes) using chiˆ2 feature selection technique along with BoW and TF-IDF feature representation techniques using different combinations of n-grams. This proposed model achieved an accuracy of 76.96% while using linear SVM with TF-IDF feature representation technique.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

feature selectionLanguage Identification

Methods 이 논문이 사용한 방법론

Feature Selection Feature selection, also known as variable selection, attribute selection or variable subset selection, is the process of selecting a subset of relevant features (variables,…

Similar Papers 제목 키워드 기반

Measles Rash Identification Using Residual Deep Convolutional Neural Network

2020-05-18 · Kimberly Glock, Charlie Napier, Andre Louie, Todd Gary 외

Measles is extremely contagious and is one of the leading causes of vaccine-preventable illness and death in developing countries, claiming more than 100,000 lives each year. Measles was declared eliminated in the US in …

SpecificityUnsupervised Pre-training

Hitachi at SemEval-2020 Task 12: Offensive Language Identification with Noisy Labels using Statistical Sampling and Post-Processing

2020-05-01 · SEMEVAL 2020 · Manikandan Ravikiran, Amin Ekant Muljibhai, Toshinori Miyoshi, Hiroaki Ozaki 외

In this paper, we present our participation in SemEval-2020 Task-12 Subtask-A (English Language) which focuses on offensive language identification from noisy labels. To this end, we developed a hybrid system with the BE…

Language IdentificationPosition

Galileo at SemEval-2020 Task 12: Multi-lingual Learning for Offensive Language Identification using Pre-trained Language Models

2020-10-07 · SEMEVAL 2020 · Shuohuan Wang, Jiaxiang Liu, Xuan Ouyang, Yu Sun

This paper describes Galileo's performance in SemEval-2020 Task 12 on detecting and categorizing offensive language in social media. For Offensive Language Identification, we proposed a multi-lingual method using Pre-tra…

AllKnowledge DistillationLanguage IdentificationXLM-R

JUNLP@DravidianLangTech-EACL2021: Offensive Language Identification in Dravidian Langauges

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Avishek Garain, Atanu Mandal, Sudip Kumar Naskar

Offensive language identification has been an active area of research in natural language processing. With the emergence of multiple social media platforms offensive language identification has emerged as a need of the h…

Language Identification

IR3218-UI at SemEval-2020 Task 12: Emoji Effects on Offensive Language IdentifiCation

2020-12-01 · SEMEVAL 2020 · Sandy Kurniawan, Indra Budi, Muhammad Okky Ibrohim

In this paper, we present our approach and the results of our participation in OffensEval 2020. There are three sub-tasks in OffensEval 2020 namely offensive language identification (sub-task A), automatic categorization…

Language Identification