paper-with-me

Papers

A Dataset for Detecting Humor in Telugu Social Media Text

2022-05-01 · DravidianLangTech (ACL) 2022 5 · Sriphani Bellamkonda, Maithili Lohakare, Shaswat Patel

Increased use of online social media sites has given rise to tremendous amounts of user generated data. Social media sites have become a platform where users express and voice their opinions in a real-time environment. Social media sites such as Twitter limit the number of characters used to express a thought in a tweet, leading to increased use of creative, humorous and confusing language in order to convey the message. Due to this, automatic humor detection has become a difficult task, especially for low-resource languages such as the Dravidian languages. Humor detection has been a well studied area for resource rich languages due to the availability of rich and accurate data. In this paper, we have attempted to solve this issue by working on low-resource languages, such as, Telugu, a Dravidian language, by collecting and annotating Telugu tweets and performing automatic humor detection on the collected data. We experimented on the corpus using various transformer models such as Multilingual BERT, Multilingual DistillBERT and XLM-RoBERTa to establish a baseline classification system. We concluded that XLM-RoBERTa was the best-performing model and it achieved an F1-score of 0.82 with 81.5% accuracy.

📄 PDF Abstract BibTeX

Code (1)

shaswa123/telugu_humour_dataset 공식 구현

Tasks

Humor Detection

Similar Papers 제목 키워드 기반

Humor Detection in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System

2018-06-14 · LREC 2018 5 · Ankush Khandelwal, Sahil Swami, Syed S. Akhtar, Manish Shrivastava

The tremendous amount of user generated data through social networking sites led to the gaining popularity of automatic text classification in the field of computational linguistics over the past decade. Within this doma…

General ClassificationHumor DetectionSentencetext-classification+1

Creating and Evaluating Code-Mixed Nepali-English and Telugu-English Datasets for Abusive Language Detection Using Traditional and Deep Learning Models

2025-04-23 · Manish Pandey, Nageshwar Prasad Yadav, Mokshada Adduru, Sawan Rai

With the growing presence of multilingual users on social media, detecting abusive language in code-mixed text has become increasingly challenging. Code-mixed communication, where users seamlessly switch between English …

Abuse DetectionAbusive Language

Developing Conversational Data and Detection of Conversational Humor in Telugu

2021-11-01 · CODI 2021 11 · Vaishnavi Pamulapati, Radhika Mamidi

In the field of humor research, there has been a recent surge of interest in the sub-domain of Conversational Humor (CH). This study has two main objectives. (a) develop a conversational (humorous and non-humorous) datas…

Transfer LearningWord Embeddings

Deceptive Humor: A Synthetic Multilingual Benchmark Dataset for Bridging Fabricated Claims with Humorous Content

2025-03-20 · Sai Kartheek Reddy Kasu, Shankar Biradar, Sunil Saumya

This paper presents the Deceptive Humor Dataset (DHD), a novel resource for studying humor derived from fabricated claims and misinformation. In an era of rampant misinformation, understanding how humor intertwines with …

Humor DetectionMisinformation

CHoRaL: Collecting Humor Reaction Labels from Millions of Social Media Users

2021-11-01 · EMNLP 2021 11 · Zixiaofan Yang, Shayan Hooshmand, Julia Hirschberg

Humor detection has gained attention in recent years due to the desire to understand user-generated content with figurative language. However, substantial individual and cultural differences in humor perception make it v…

Humor Detection