paper-with-me

Papers

Spectro-Temporal Modulation Representation Framework for Human-Imitated Speech Detection

2026-04-25 · Khalid Zaman, Masashi Unoki arxiv

Human-imitated speech poses a greater challenge than AI-generated speech for both human listeners and automatic detection systems. Unlike AI-generated speech, which often contains artifacts, over-smoothed spectra, or robotic cues, imitated speech is produced naturally by humans, thereby preserving a higher degree of naturalness that makes imitation-based speech forgery significantly more challenging to detect using conventional acoustic or cepstral features. To overcome this challenge, this study proposes an auditory perception-based Spectro-Temporal Modulation (STM) representation framework for human-imitated speech detection. The STM representations are derived from two cochlear filterbank models: the Gammatone Filterbank (GTFB), which simulates frequency selectivity and can be regarded as a first approximation of cochlear filtering, and the Gammachirp Filterbank (GCFB), which further models both frequency selectivity and level-dependent asymmetry. These STM representations jointly capture temporal and spectral fluctuations in speech signals, corresponding to changes over time in the spectrogram and variations along the frequency axis related to human auditory perception. We also introduce a Segmental-STM representation to analyze short-term modulation patterns across overlapping time windows, enabling high-resolution modeling of temporal speech variations. Experimental results show that STM representations are effective for human-imitated speech detection, achieving accuracy levels close to those of human listeners. In addition, Segmental-STM representations are more effective, surpassing human perceptual performance. The findings demonstrate that perceptually inspired spectro-temporal modeling is promising for detecting imitation-based speech attacks and improving voice authentication robustness.

📄 PDF Abstract BibTeX arXiv:2604.23241

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Modulation spectral features for speech emotion recognition using deep neural networks

2023-01-14 · Premjeet Singh, Md Sahidullah, Goutam Saha

This work explores the use of constant-Q transform based modulation spectral features (CQT-MSF) for speech emotion recognition (SER). The human perception and analysis of sound comprise of two important cognitive parts: …

Emotion RecognitionSpeech Emotion Recognition

Learning spectro-temporal representations of complex sounds with parameterized neural networks

2021-03-12 · Rachid Riad, Julien Karadayi, Anne-Catherine Bachoud-Lévi, Emmanuel Dupoux

Deep Learning models have become potential candidates for auditory neuroscience research, thanks to their recent successes on a variety of auditory tasks. Yet, these models often lack interpretability to fully understand…

Action DetectionActivity DetectionSound ClassificationSpeaker Verification

Spectrotemporal Modulation: Efficient and Interpretable Feature Representation for Classifying Speech, Music, and Environmental Sounds

2025-05-29 · Andrew Chang, Yike Li, Iran R. Roman, David Poeppel

Audio DNNs have demonstrated impressive performance on various machine listening tasks; however, most of their representations are computationally costly and uninterpretable, leaving room for optimization. Here, we propo…

Audio Classification

On combining acoustic and modulation spectrograms in an attention LSTM-based system for speech intelligibility level classification

2024-02-05 · Ascensión Gallardo-Antolín, Juan M. Montero

Speech intelligibility can be affected by multiple factors, such as noisy environments, channel distortions or physiological issues. In this work, we deal with the problem of automatic prediction of the speech intelligib…

Differentiable Time-Frequency Scattering on GPU

2022-04-18 · John Muradeli, Cyrus Vahidi, Changhong Wang, Han Han 외

Joint time-frequency scattering (JTFS) is a convolutional operator in the time-frequency domain which extracts spectrotemporal modulations at various rates and scales. It offers an idealized model of spectrotemporal rece…

Audio GenerationCPUGPUResynthesis