paper-with-me

Papers

ProSarc: Prosody-Aware Sarcasm Recognition Framework via Temporal Prosodic Incongruity

2026-06-04 · Prathamjyot Singh, Ashima Sood, Sahil Sharma, Jasmeet Singh arxiv

We present ProSarc, an audio-only framework that detects sarcasm by modelling temporal prosodic incongruity, that is, the mismatch between local prosodic dynamics and the utterance-level emotional baseline. Dual encoding paths, a Global Emotion Encoder and a Temporal Prosody Encoder (BiLSTM + multi-head attention), feed a Prosodic Incongruity Analyzer that produces a scalar incongruity score for classification. Monte Carlo dropout provides uncertainty estimates, and an attention-based mechanism localises sarcastic onset without frame-level labels. ProSarc outperforms prior audio-only methods on MUStARD++ (F1=75.3) and generalises to spontaneous (PodSarc, F1=62.9) and cross-lingual speech (MuSaG, F1=65.6). Ten-run validation confirms the contribution of incongruity modelling (Wilcoxon p=0.002, Cohen's d=1.51). Human evaluation shows that model uncertainty tracks perceptual ambiguity and predicted onsets align with human-annotated temporal windows.

📄 PDF Abstract BibTeX arXiv:2606.06168

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Modeling Sarcastic Speech: Semantic and Prosodic Cues in a Speech Synthesis Framework

2025-10-08 · Zhu Li, Yuqing Zhang, Xiyuan Gao, Shekhar Nayak 외 arxiv

Sarcasm is a pragmatic phenomenon in which speakers convey meanings that diverge from literal content, relying on an interaction between semantics and prosodic expression. However, how these cues jointly contribute to th…

Speech Synthesis

Integrating Feedback Loss from Bi-modal Sarcasm Detector for Sarcastic Speech Synthesis

2025-08-18 · Zhu Li, Yuqing Zhang, Xiyuan Gao, Devraj Raghuvanshi 외 arxiv

Sarcastic speech synthesis, which involves generating speech that effectively conveys sarcasm, is essential for enhancing natural interactions in applications such as entertainment and human-computer interaction. However…

Sarcasm DetectionTransfer LearningSpeech Synthesis

CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection

2026-09-15 · Qiyang Sun, Xudong Li, Yupei Li, Jiabin Xue 외 arxiv

Spoken sarcasm detectors may exploit lexical content, prosody, or their interaction, yet conventional evaluation cannot reveal which cues drive their predictions. We introduce CLASH (Controlled Lexical-Acoustic Separatio…

Self-Supervised LearningSarcasm Detection

What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels

2025-12-18 · Aditya Yadavalli, Tiago Pimentel, Tamar I Regev, Ethan Wilcox 외 arxiv

Prosody -- the melody of speech -- conveys critical information often not captured by the words or text of a message. In this paper, we propose an information-theoretic approach to quantify how much information is expres…

Seeing Sarcasm Through Different Eyes: Analyzing Multimodal Sarcasm Perception in Large Vision-Language Models

2025-03-15 · Junjie Chen, Xuyang Liu, Subin Huang, Linfeng Zhang 외

With the advent of large vision-language models (LVLMs) demonstrating increasingly human-like abilities, a pivotal question emerges: do different LVLMs interpret multimodal sarcasm differently, and can a single model gra…