paper-with-me

홈 › Papers

Investigating Causal Cues: Strengthening Spoofed Audio Detection with Human-Discernible Linguistic Features

2024-09-09 · Zahra Khanjani, Tolulope Ale, Jianwu Wang, Lavon Davis, Christine Mallinson, Vandana P. Janeja

Several types of spoofed audio, such as mimicry, replay attacks, and deepfakes, have created societal challenges to information integrity. Recently, researchers have worked with sociolinguistics experts to label spoofed audio samples with Expert Defined Linguistic Features (EDLFs) that can be discerned by the human ear: pitch, pause, word-initial and word-final release bursts of consonant stops, audible intake or outtake of breath, and overall audio quality. It is established that there is an improvement in several deepfake detection algorithms when they augmented the traditional and common features of audio data with these EDLFs. In this paper, using a hybrid dataset comprised of multiple types of spoofed audio augmented with sociolinguistic annotations, we investigate causal discovery and inferences between the discernible linguistic features and the label in the audio clips, comparing the findings of the causal models with the expert ground truth validation labeling process. Our findings suggest that the causal models indicate the utility of incorporating linguistic features to help discern spoofed audio, as well as the overall need and opportunity to incorporate human knowledge into models and techniques for strengthening AI models. The causal discovery and inference can be used as a foundation of training humans to discern spoofed audio as well as automating EDLFs labeling for the purpose of performance improvement of the common AI-based spoofed audio detectors.

📄 PDF Abstract BibTeX arXiv:2409.06033

Code (0)

등록된 구현이 없습니다.

Tasks

Causal DiscoveryDeepFake DetectionFace Swapping

Similar Papers 제목 키워드 기반

Linguistically Augmented Audio Speech Data (LinguAS)

2026-06-08 · Ashley R. Keaton, Zahra Khanjani, Christine Mallinson, Vandana P. Janeja arxiv

Maliciously-created fake speech, including deepfaked and spoofed audio, is proliferating at an alarming rate, and detection models are racing to stay ahead of the curve. Yet, most detection models are trained to make inf…

How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?

2024-06-04 · Tianchi Liu, Lin Zhang, Rohan Kumar Das, Yi Ma 외

Partially manipulating a sentence can greatly change its meaning. Recent work shows that countermeasures (CMs) trained on partially spoofed audio can effectively detect such spoofing. However, the current understanding o…

Decision MakingSentence

Waveform Boundary Detection for Partially Spoofed Audio

2022-11-01 · Zexin Cai, Weiqing Wang, Ming Li

The present paper proposes a waveform boundary detection system for audio spoofing attacks containing partially manipulated segments. Partially spoofed/fake audio, where part of the utterance is replaced, either with syn…

Boundary Detection

An Initial Investigation for Detecting Partially Spoofed Audio

2021-04-06 · Lin Zhang, Xin Wang, Erica Cooper, Junichi Yamagishi 외

All existing databases of spoofed speech contain attack data that is spoofed in its entirety. In practice, it is entirely plausible that successful attacks can be mounted with utterances that are only partially spoofed. …

Voice Anti-spoofing

Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues

2024-02-04 · Tianxiang Chen, Zhentao Tan, Tao Gong, Qi Chu 외

How to effectively interact audio with vision has garnered considerable interest within the multi-modality research field. Recently, a novel audio-visual segmentation (AVS) task has been proposed, aiming to segment the s…

DecoderRepresentation Learning