paper-with-me

홈 › Papers

MMT: A Multilingual and Multi-Topic Indian Social Media Dataset

2023-04-02 · Dwip Dalal, Vivek Srivastava, Mayank Singh

Social media plays a significant role in cross-cultural communication. A vast amount of this occurs in code-mixed and multilingual form, posing a significant challenge to Natural Language Processing (NLP) tools for processing such information, like language identification, topic modeling, and named-entity recognition. To address this, we introduce a large-scale multilingual, and multi-topic dataset (MMT) collected from Twitter (1.7 million Tweets), encompassing 13 coarse-grained and 63 fine-grained topics in the Indian context. We further annotate a subset of 5,346 tweets from the MMT dataset with various Indian languages and their code-mixed counterparts. Also, we demonstrate that the currently existing tools fail to capture the linguistic diversity in MMT on two downstream tasks, i.e., topic modeling and language identification. To facilitate future research, we will make the anonymized and annotated dataset available in the public domain.

📄 PDF Abstract BibTeX arXiv:2304.00634

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityLanguage Identificationnamed-entity-recognitionNamed Entity Recognition

Methods 이 논문이 사용한 방법론

fail 설명 없음

Similar Papers 제목 키워드 기반

Recurrent Neural Network based Part-of-Speech Tagger for Code-Mixed Social Media Text

2016-11-15 · Raj Nath Patel, Prakash B. Pimpale, M Sasikumar

This paper describes Centre for Development of Advanced Computing's (CDACM) submission to the shared task-'Tool Contest on POS tagging for Code-Mixed Indian Social Media (Facebook, Twitter, and Whatsapp) Text', collocate…

Language ModelingLanguage ModellingPOSPOS Tagging+1

Lost in Translation, Found in Spans: Identifying Claims in Multilingual Social Media

2023-10-27 · Shubham Mittal, Megha Sundriyal, Preslav Nakov

Claim span identification (CSI) is an important step in fact-checking pipelines, aiming to identify text segments that contain a checkworthy claim or assertion in a social media post. Despite its importance to journalist…

Cross-Lingual TransferFact CheckingXLM-R

Multilingual Topic Classification in X: Dataset and Analysis

2024-10-04 · Dimosthenis Antypas, Asahi Ushio, Francesco Barbieri, Jose Camacho-Collados

In the dynamic realm of social media, diverse topics are discussed daily, transcending linguistic boundaries. However, the complexities of understanding and categorising this content across various languages remain an im…

ClassificationDiversityTopic Classification

muBoost: An Effective Method for Solving Indic Multilingual Text Classification Problem

2022-06-21 · Manish Pathak, Aditya Jain

Text Classification is an integral part of many Natural Language Processing tasks such as sarcasm detection, sentiment analysis and many more such applications. Many e-commerce websites, social-media/entertainment platfo…

Multilingual text classificationSarcasm DetectionSentiment Analysistext-classification+1

Can Multilingual Transformers Fight the COVID-19 Infodemic?

2021-09-01 · RANLP 2021 9 · Lasitha Uyangodage, Tharindu Ranasinghe, Hansi Hettiarachchi

The massive spread of false information on social media has become a global risk especially in a global pandemic situation like COVID-19. False information detection has thus become a surging research topic in recent mon…

BIG-bench Machine Learning