paper-with-me

Papers

Preparing Bengali-English Code-Mixed Corpus for Sentiment Analysis of Indian Languages

2018-03-11 · Soumil Mandal, Sainik Kumar Mahata, Dipankar Das

Analysis of informative contents and sentiments of social users has been attempted quite intensively in the recent past. Most of the systems are usable only for monolingual data and fails or gives poor results when used on data with code-mixing property. To gather attention and encourage researchers to work on this crisis, we prepared gold standard Bengali-English code-mixed data with language and polarity tag for sentiment analysis purposes. In this paper, we discuss the systems we prepared to collect and filter raw Twitter data. In order to reduce manual work while annotation, hybrid systems combining rule based and supervised models were developed for both language and sentiment tagging. The final corpus was annotated by a group of annotators following a few guidelines. The gold standard corpus thus obtained has impressive inter-annotator agreement obtained in terms of Kappa values. Various metrics like Code-Mixed Index (CMI), Code-Mixed Factor (CF) along with various aspects (language and emotion) also qualitatively polled the code-mixed and sentiment properties of the corpus.

📄 PDF Abstract BibTeX arXiv:1803.04000

Code (0)

등록된 구현이 없습니다.

Tasks

Sentiment AnalysisTAG

Similar Papers 제목 키워드 기반

Development of POS tagger for English-Bengali Code-Mixed data

2020-07-29 · ICON 2019 12 · Tathagata Raha, Sainik Kumar Mahata, Dipankar Das, Sivaji Bandyopadhyay

Code-mixed texts are widespread nowadays due to the advent of social media. Since these texts combine two languages to formulate a sentence, it gives rise to various research problems related to Natural Language Processi…

POSSentenceTAG

Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation

2021-07-17 · ACL ARR November 2021 11 · Anonymous

The widespread online communication in a modern multilingual world has provided opportunities to blend more than one language (aka. code-mixed language) in a single utterance. This has resulted a formidable challenge for…

Machine TranslationSentenceSynthetic Data GenerationTranslation

Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation

2024-03-25 · Kartik Kartik, Sanjana Soni, Anoop Kunchukuttan, Tanmoy Chakraborty 외

The widespread online communication in a modern multilingual world has provided opportunities to blend more than one language (aka code-mixed language) in a single utterance. This has resulted a formidable challenge for …

Machine TranslationSentenceSynthetic Data GenerationTranslation

BnSentMix: A Diverse Bengali-English Code-Mixed Dataset for Sentiment Analysis

2024-08-16 · Sadia Alam, Md Farhan Ishmam, Navid Hasin Alvee, Md Shahnewaz Siddique 외

The widespread availability of code-mixed data can provide valuable insights into low-resource languages like Bengali, which have limited datasets. Sentiment analysis has been a fundamental text classification task acros…

DiversitySentiment AnalysisSentiment Classificationtext-classification+1

JU_KS@SAIL_CodeMixed-2017: Sentiment Analysis for Indian Code Mixed Social Media Texts

2018-02-15 · Kamal Sarkar

This paper reports about our work in the NLP Tool Contest @ICON-2017, shared task on Sentiment Analysis for Indian Languages (SAIL) (code mixed). To implement our system, we have used a machine learning algo-rithm called…

PositionSentiment Analysis