paper-with-me

Papers

ToxicTone: A Mandarin Audio Dataset Annotated for Toxicity and Toxic Utterance Tonality

2025-05-21 · Yu-Xiang Luo, Yi-Cheng Lin, Ming-To Chuang, Jia-Hung Chen, I-Ning Tsai, Pei Xing Kiew, Yueh-Hsuan Huang, Chien-Feng Liu, Yu-Chen Chen, Bo-Han Feng, Wenze Ren, Hung-Yi Lee

Despite extensive research on toxic speech detection in text, a critical gap remains in handling spoken Mandarin audio. The lack of annotated datasets that capture the unique prosodic cues and culturally specific expressions in Mandarin leaves spoken toxicity underexplored. To address this, we introduce ToxicTone -- the largest public dataset of its kind -- featuring detailed annotations that distinguish both forms of toxicity (e.g., profanity, bullying) and sources of toxicity (e.g., anger, sarcasm, dismissiveness). Our data, sourced from diverse real-world audio and organized into 13 topical categories, mirrors authentic communication scenarios. We also propose a multimodal detection framework that integrates acoustic, linguistic, and emotional features using state-of-the-art speech and emotion encoders. Extensive experiments show our approach outperforms text-only and baseline models, underscoring the essential role of speech-specific cues in revealing hidden toxic expressions.

📄 PDF Abstract BibTeX arXiv:2505.15773

Code (1)

YuXiangLo/ToxicTone 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Audio Caption: Listen and Tell

2019-02-25 · Mengyue Wu, Heinrich Dinkel, Kai Yu

Increasing amount of research has shed light on machine perception of audio events, most of which concerns detection and classification tasks. However, human-like perception of audio scenes involves not only detecting an…

DecoderGeneral Classification

Audio Caption in a Car Setting with a Sentence-Level Loss

2019-05-31 · Xuenan Xu, Heinrich Dinkel, Mengyue Wu, Kai Yu

Captioning has attracted much attention in image and video understanding while a small amount of work examines audio captioning. This paper contributes a Mandarin-annotated dataset for audio captioning within a car scene…

Audio captioningDecoderSemantic SimilaritySemantic Textual Similarity+4

Annotating a corpus of human interaction with prosodic profiles --- focusing on Mandarin repair/disfluency

2012-05-01 · LREC 2012 5 · Helen Kai-yun Chen

This study describes the construction of a manually annotated speech corpus that focuses on the sound profiles of repair/disfluency in Mandarin conversational interaction. Specifically, the paper focuses on how the tag s…

TAG

MuTox: Universal MUltilingual Audio-based TOXicity Dataset and Zero-shot Detector

2024-01-10 · Marta R. Costa-jussà, Mariano Coria Meglioli, Pierre Andrews, David Dale 외

Research in toxicity detection in natural language processing for the speech modality (audio-based) is quite limited, particularly for languages other than English. To address these limitations and lay the groundwork for…

FMFCC-A: A Challenging Mandarin Dataset for Synthetic Speech Detection

2021-10-18 · Zhenyu Zhang, Yewei Gu, Xiaowei Yi, Xianfeng Zhao

As increasing development of text-to-speech (TTS) and voice conversion (VC) technologies, the detection of synthetic speech has been suffered dramatically. In order to promote the development of synthetic speech detectio…

Speech SynthesisSynthetic Speech Detectiontext-to-speechText to Speech+1