Frequency Domain-Based Detection of Generated Audio
Attackers may manipulate audio with the intent of presenting falsified reports, changing an opinion of a public figure, and winning influence and power. The prevalence of inauthentic multimedia continues to rise, so it is imperative to develop a set of tools that determines the legitimacy of media. We present a method that analyzes audio signals to determine whether they contain real human voices or fake human voices (i.e., voices generated by neural acoustic and waveform models). Instead of analyzing the audio signals directly, the proposed approach converts the audio signals into spectrogram images displaying frequency, intensity, and temporal content and evaluates them with a Convolutional Neural Network (CNN). Trained on both genuine human voice signals and synthesized voice signals, we show our approach achieves high accuracy on this classification task.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Pitch Contour Exploration Across Audio Domains: A Vision-Based Transfer Learning Approach
This study examines pitch contours as a unifying semantic construct prevalent across various audio domains including music, speech, bioacoustics, and everyday sounds. Analyzing pitch contours offers insights into the uni…
object-detectionObject DetectionTransfer LearningFrom Vision to Sound: Advancing Audio Anomaly Detection with Vision-Based Algorithms
Recent advances in Visual Anomaly Detection (VAD) have introduced sophisticated algorithms leveraging embeddings generated by pre-trained feature extractors. Inspired by these developments, we investigate the adaptation …
Anomaly DetectionAudios Don't Lie: Multi-Frequency Channel Attention Mechanism for Audio Deepfake Detection
With the rapid development of artificial intelligence technology, the application of deepfake technology in the audio field has gradually increased, resulting in a wide range of security risks. Especially in the financia…
Audio Deepfake DetectionDeepFake DetectionFace SwappingA Multi-Domain Feature Fusion Framework for Generalizable Deepfake Detection Across Different Generators
Deepfakes are artificially generated images, audio, or videos that threaten privacy, security, and information integrity. Detecting such content is crucial for countering disinformation, as the latest models generate hig…
DeepFake DetectionData AugmentationTime-weighted Frequency Domain Audio Representation with GMM Estimator for Anomalous Sound Detection
Although deep learning is the mainstream method in unsupervised anomalous sound detection, Gaussian Mixture Model (GMM) with statistical audio frequency representation as input can achieve comparable results with much lo…