paper-with-me

홈 › Papers

A New Approach to Voice Authenticity

2024-02-09 · Nicolas M. Müller, Piotr Kawa, Shen Hu, Matthias Neu, Jennifer Williams, Philip Sperl, Konstantin Böttinger

Voice faking, driven primarily by recent advances in text-to-speech (TTS) synthesis technology, poses significant societal challenges. Currently, the prevailing assumption is that unaltered human speech can be considered genuine, while fake speech comes from TTS synthesis. We argue that this binary distinction is oversimplified. For instance, altered playback speeds can be used for malicious purposes, like in the 'Drunken Nancy Pelosi' incident. Similarly, editing of audio clips can be done ethically, e.g., for brevity or summarization in news reporting or podcasts, but editing can also create misleading narratives. In this paper, we propose a conceptual shift away from the binary paradigm of audio being either 'fake' or 'real'. Instead, our focus is on pinpointing 'voice edits', which encompass traditional modifications like filters and cuts, as well as TTS synthesis and VC systems. We delineate 6 categories and curate a new challenge dataset rooted in the M-AILABS corpus, for which we present baseline detection systems. And most importantly, we argue that merely categorizing audio as fake or real is a dangerous over-simplification that will fail to move the field of speech technology forward.

📄 PDF Abstract BibTeX arXiv:2402.06304

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to Speech

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

What Counts as Real? Speech Restoration and Voice Quality Conversion Pose New Challenges to Deepfake Detection

2026-03-14 · Shree Harsha Bokkahalli Satish, Harm Lameris, Joakim Gustafson, Éva Székely arxiv

Audio anti-spoofing systems are typically trained to assign one authenticity label to an entire speech utterance. This formulation becomes under-specified for transformations where the underlying speaker identity and lin…

DeepFake Detection

FSD: An Initial Chinese Dataset for Fake Song Detection

2023-09-05 · Yuankun Xie, Jingjing Zhou, Xiaolin Lu, Zhenghao Jiang 외

Singing voice synthesis and singing voice conversion have significantly advanced, revolutionizing musical experiences. However, the rise of "Deepfake Songs" generated by these technologies raises concerns about authentic…

Audio Deepfake DetectionDeepFake DetectionFace SwappingFake Song Detection+2

Singing Voice Graph Modeling for SingFake Detection

2024-06-05 · Xuanjun Chen, Haibin Wu, Jyh-Shing Roger Jang, Hung-Yi Lee

Detecting singing voice deepfakes, or SingFake, involves determining the authenticity and copyright of a singing voice. Existing models for speech deepfake detection have struggled to adapt to unseen attacks in this uniq…

DeepFake DetectionFace SwappingRhythm

Does Listening Matter? Backchanneling and Nodding in AI Clone

2026-08-20 · Koji Inoue, Kazushi Kato, Tatsuya Kawahara, Shunichi Kasahara arxiv

AI clones that imitate a specific person typically reproduce what the person says and how they sound, but not how they listen. We investigate whether adding multimodal listening behaviors gives such a clone more presence…

Easy, Interpretable, Effective: openSMILE for voice deepfake detection

2024-08-28 · Octavian Pascu, Dan Oneata, Horia Cucu, Nicolas M. Müller

In this paper, we demonstrate that attacks in the latest ASVspoof5 dataset -- a de facto standard in the field of voice authenticity and deepfake detection -- can be identified with surprising accuracy using a small subs…

DeepFake DetectionFace Swappingtext-to-speechText to Speech+1