paper-with-me

홈 › Papers

EmoMix-3L: A Code-Mixed Dataset for Bangla-English-Hindi Emotion Detection

2024-05-11 · Nishat Raihan, Dhiman Goswami, Antara Mahmud, Antonios Anastasopoulos, Marcos Zampieri

Code-mixing is a well-studied linguistic phenomenon that occurs when two or more languages are mixed in text or speech. Several studies have been conducted on building datasets and performing downstream NLP tasks on code-mixed data. Although it is not uncommon to observe code-mixing of three or more languages, most available datasets in this domain contain code-mixed data from only two languages. In this paper, we introduce EmoMix-3L, a novel multi-label emotion detection dataset containing code-mixed data from three different languages. We experiment with several models on EmoMix-3L and we report that MuRIL outperforms other models on this dataset.

📄 PDF Abstract BibTeX arXiv:2405.06922

Code (1)

GoswamiDhiman/EmoMix-3L 공식 구현

Similar Papers 제목 키워드 기반

Word-level Language Identification Using Subword Embeddings for Code-mixed Bangla-English Social Media Data

2022-06-01 · DCLRL (LREC) 2022 6 · Aparna Dutta

This paper reports work on building a word-level language identification (LID) model for code-mixed Bangla-English social media data using subword embeddings, with an ultimate goal of using this LID module as the first s…

Language IdentificationPOS

SentMix-3L: A Bangla-English-Hindi Code-Mixed Dataset for Sentiment Analysis

2023-10-27 · Md Nishat Raihan, Dhiman Goswami, Antara Mahmud, Antonios Anastasopoulos 외

Code-mixing is a well-studied linguistic phenomenon when two or more languages are mixed in text or speech. Several datasets have been build with the goal of training computational models for code-mixing. Although it is …

Sentiment Analysis

MixSarc: A Bangla-English Code-Mixed Corpus for Implicit Meaning Identification

2026-02-25 · Kazi Samin Yasar Alam, Md Tanbir Chowdhury, Tamim Ahmed, Ajwad Abrar 외 arxiv

Bangla-English code-mixing is widespread across South Asian social media, yet resources for implicit meaning identification in this setting remain scarce. Existing sentiment and sarcasm models largely focus on monolingua…

Humor Detection

Mixed-Distil-BERT: Code-mixed Language Modeling for Bangla, English, and Hindi

2023-09-19 · Md Nishat Raihan, Dhiman Goswami, Antara Mahmud

One of the most popular downstream tasks in the field of Natural Language Processing is text classification. Text classification tasks have become more daunting when the texts are code-mixed. Though they are not exposed …

Language ModelingLanguage Modellingtext-classificationText Classification+1

OffMix-3L: A Novel Code-Mixed Dataset in Bangla-English-Hindi for Offensive Language Identification

2023-10-27 · Dhiman Goswami, Md Nishat Raihan, Antara Mahmud, Antonios Anastasopoulos 외

Code-mixing is a well-studied linguistic phenomenon when two or more languages are mixed in text or speech. Several works have been conducted on building datasets and performing downstream NLP tasks on code-mixed data. A…

Language Identification