paper-with-me

Papers

COMMENTATOR: A Code-mixed Multilingual Text Annotation Framework

2024-08-06 · Rajvee Sheth, Shubh Nisar, Heenaben Prajapati, Himanshu Beniwal, Mayank Singh

As the NLP community increasingly addresses challenges associated with multilingualism, robust annotation tools are essential to handle multilingual datasets efficiently. In this paper, we introduce a code-mixed multilingual text annotation framework, COMMENTATOR, specifically designed for annotating code-mixed text. The tool demonstrates its effectiveness in token-level and sentence-level language annotation tasks for Hinglish text. We perform robust qualitative human-based evaluations to showcase COMMENTATOR led to 5x faster annotations than the best baseline. Our code is publicly available at \url{https://github.com/lingo-iitgn/commentator}. The demonstration video is available at \url{https://bit.ly/commentator_video}.

📄 PDF Abstract BibTeX arXiv:2408.03125

Code (1)

lingo-iitgn/commentator 공식 구현

Tasks

Sentencetext annotation

Similar Papers 제목 키워드 기반

COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing

2025-03-27 · Rajvee Sheth, Himanshu Beniwal, Mayank Singh

The rapid growth of digital communication has driven the widespread use of code-mixing, particularly Hindi-English, in multilingual communities. Existing datasets often focus on romanized text, have limited scope, or rel…

Language Identificationnamed-entity-recognitionNamed Entity RecognitionPart-Of-Speech Tagging

cs@DravidianLangTech-EACL2021: Offensive Language Identification Based On Multilingual BERT Model

2021-04-01 · EACL (DravidianLangTech) 2021 4 · Shi Chen, Bing Kong

This paper introduces the related content of the task “Offensive Language Identification in Dravidian LANGUAGES-EACL 2021”. The task requires us to classify Dravidian languages collected from social media into Not-Offens…

Language Identificationtext-classificationText Classification

Advancing Sentiment Analysis in Tamil-English Code-Mixed Texts: Challenges and Transformer-Based Solutions

2025-03-30 · Mikhail Krasitskii, Olga Kolesnikova, Liliana Chanona Hernandez, Grigori Sidorov 외

The sentiment analysis task in Tamil-English code-mixed texts has been explored using advanced transformer-based models. Challenges from grammatical inconsistencies, orthographic variations, and phonetic ambiguities have…

Data AugmentationSentiment AnalysisSentiment Classification

FOOCTTS: Generating Arabic Speech with Acoustic Environment for Football Commentator

2023-06-07 · Massa Baali, Ahmed Ali

This paper presents FOOCTTS, an automatic pipeline for a football commentator that generates speech with background crowd noise. The application gets the text from the user, applies text pre-processing such as vowelizati…

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

EM2LDL: A Multilingual Speech Corpus for Mixed Emotion Recognition through Label Distribution Learning

2025-11-25 · Xingfeng Li, Xiaohan Shi, Junjie Li, Yongwei Li 외 arxiv

This study introduces EM2LDL, a novel multilingual speech corpus designed to advance mixed emotion recognition through label distribution learning. Addressing the limitations of predominantly monolingual and single-label…

Self-Supervised LearningEmotion Recognition