MMT: A Multilingual and Multi-Topic Indian Social Media Dataset
Social media plays a significant role in cross-cultural communication. A vast amount of this occurs in code-mixed and multilingual form, posing a significant challenge to Natural Language Processing (NLP) tools for processing such information, like language identification, topic modeling, and named-entity recognition. To address this, we introduce a large-scale multilingual, and multi-topic dataset (MMT) collected from Twitter (1.7 million Tweets), encompassing 13 coarse-grained and 63 fine-grained topics in the Indian context. We further annotate a subset of 5,346 tweets from the MMT dataset with various Indian languages and their code-mixed counterparts. Also, we demonstrate that the currently existing tools fail to capture the linguistic diversity in MMT on two downstream tasks, i.e., topic modeling and language identification. To facilitate future research, we will make the anonymized and annotated dataset available in the public domain.
Code (0)
등록된 구현이 없습니다.
Tasks
DiversityLanguage Identificationnamed-entity-recognitionNamed Entity RecognitionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Recurrent Neural Network based Part-of-Speech Tagger for Code-Mixed Social Media Text
This paper describes Centre for Development of Advanced Computing's (CDACM) submission to the shared task-'Tool Contest on POS tagging for Code-Mixed Indian Social Media (Facebook, Twitter, and Whatsapp) Text', collocate…
Language ModelingLanguage ModellingPOSPOS Tagging+1Lost in Translation, Found in Spans: Identifying Claims in Multilingual Social Media
Claim span identification (CSI) is an important step in fact-checking pipelines, aiming to identify text segments that contain a checkworthy claim or assertion in a social media post. Despite its importance to journalist…
Cross-Lingual TransferFact CheckingXLM-RMultilingual Topic Classification in X: Dataset and Analysis
In the dynamic realm of social media, diverse topics are discussed daily, transcending linguistic boundaries. However, the complexities of understanding and categorising this content across various languages remain an im…
ClassificationDiversityTopic ClassificationmuBoost: An Effective Method for Solving Indic Multilingual Text Classification Problem
Text Classification is an integral part of many Natural Language Processing tasks such as sarcasm detection, sentiment analysis and many more such applications. Many e-commerce websites, social-media/entertainment platfo…
Multilingual text classificationSarcasm DetectionSentiment Analysistext-classification+1Can Multilingual Transformers Fight the COVID-19 Infodemic?
The massive spread of false information on social media has become a global risk especially in a global pandemic situation like COVID-19. False information detection has thus become a surging research topic in recent mon…
BIG-bench Machine Learning