paper-with-me

홈 › Papers

Synthetic Voice Detection and Audio Splicing Detection using SE-Res2Net-Conformer Architecture

2022-10-07 · Lei Wang, Benedict Yeoh, Jun Wah Ng

Synthetic voice and splicing audio clips have been generated to spoof Internet users and artificial intelligence (AI) technologies such as voice authentication. Existing research work treats spoofing countermeasures as a binary classification problem: bonafide vs. spoof. This paper extends the existing Res2Net by involving the recent Conformer block to further exploit the local patterns on acoustic features. Experimental results on ASVspoof 2019 database show that the proposed SE-Res2Net-Conformer architecture is able to improve the spoofing countermeasures performance for the logical access scenario. In addition, this paper also proposes to re-formulate the existing audio splicing detection problem. Instead of identifying the complete splicing segments, it is more useful to detect the boundaries of the spliced segments. Moreover, a deep learning approach can be used to solve the problem, which is different from the previous signal processing techniques.

📄 PDF Abstract BibTeX arXiv:2210.03581

Code (0)

등록된 구현이 없습니다.

Tasks

Binary Classification

Methods 이 논문이 사용한 방법론

ReLU How Do I Communicate to Expedia? How Do I Communicate to Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Live Support & Special Travel…
Residual Connection 설명 없음
Average Pooling 설명 없음
Global Average Pooling Global Average Pooling is a pooling operation designed to replace fully connected layers in classical CNNs. The idea is to generate one feature map for each corresponding…
Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…
Batch Normalization 설명 없음
1x1 Convolution A 1 x 1 Convolution is a convolution with some special properties in that it can be used for dimensionality reduction,…
Res2Net Block A Res2Net Block is an image model block that constructs hierarchical residual-like connections within one single [residual…

Similar Papers 제목 키워드 기반

Towards Unconstrained Audio Splicing Detection and Localization with Neural Networks

2022-07-29 · Denise Moussa, Germans Hirsch, Christian Riess

Freely available and easy-to-use audio editing tools make it straightforward to perform audio splicing. Convincing forgeries can be created by combining various speech samples from the same person. Detection of such spli…

Human DetectionMisinformation

Attacking Image Splicing Detection and Localization Algorithms Using Synthetic Traces

2022-11-22 · Shengbang Fang, Matthew C Stamm

Recent advances in deep learning have enabled forensics researchers to develop a new class of image splicing detection and localization algorithms. These algorithms identify spliced content by detecting localized inconsi…

Generative Adversarial Network

Gender Fairness in Audio Deepfake Detection: Performance and Disparity Analysis

2026-03-09 · Aishwarya Fursule, Shruti Kshirsagar, Anderson R. Avila arxiv

Audio deepfake detection aims to detect real human voices from those generated by Artificial Intelligence (AI) and has emerged as a significant problem in the field of voice biometrics systems. With the ever-improving qu…

Audio Deepfake Detection

Point to the Hidden: Exposing Speech Audio Splicing via Signal Pointer Nets

2023-07-11 · Denise Moussa, Germans Hirsch, Sebastian Wankerl, Christian Riess

Verifying the integrity of voice recording evidence for criminal investigations is an integral part of an audio forensic analyst's work. Here, one focus is on detecting deletion or insertion operations, so called audio s…

Defense Against Synthetic Speech: Real-Time Detection of RVC Voice Conversion Attacks

2025-12-31 · Prajwal Chinchmalatpure, Suyash Chinchmalatpure, Siddharth Chavan arxiv

Generative audio technologies now enable highly realistic voice cloning and real-time voice conversion, increasing the risk of impersonation, fraud, and misinformation in communication channels such as phone and video ca…

Voice Conversion