paper-with-me

Papers

Spectrotemporal Modulation: Efficient and Interpretable Feature Representation for Classifying Speech, Music, and Environmental Sounds

2025-05-29 · Andrew Chang, Yike Li, Iran R. Roman, David Poeppel

Audio DNNs have demonstrated impressive performance on various machine listening tasks; however, most of their representations are computationally costly and uninterpretable, leaving room for optimization. Here, we propose a novel approach centered on spectrotemporal modulation (STM) features, a signal processing method that mimics the neurophysiological representation in the human auditory cortex. The classification performance of our STM-based model, without any pretraining, is comparable to that of pretrained audio DNNs across diverse naturalistic speech, music, and environmental sounds, which are essential categories for both human cognition and machine perception. These results show that STM is an efficient and interpretable feature representation for audio classification, advancing the development of machine listening and unlocking exciting new possibilities for basic understanding of speech and auditory sciences, as well as developing audio BCI and cognitive computing.

📄 PDF Abstract BibTeX arXiv:2505.23509

Code (1)

curlsloth/MusicSpeech-STM 공식 구현 pytorch

Tasks

Audio Classification

Similar Papers 제목 키워드 기반

Differentiable Time-Frequency Scattering on GPU

2022-04-18 · John Muradeli, Cyrus Vahidi, Changhong Wang, Han Han 외

Joint time-frequency scattering (JTFS) is a convolutional operator in the time-frequency domain which extracts spectrotemporal modulations at various rates and scales. It offers an idealized model of spectrotemporal rece…

Audio GenerationCPUGPUResynthesis

Toward Knowledge-Driven Speech-Based Models of Depression: Leveraging Spectrotemporal Variations in Speech Vowels

2022-10-05 · Kexin Feng, Theodora Chaspari

Psychomotor retardation associated with depression has been linked with tangible differences in vowel production. This paper investigates a knowledge-driven machine learning (ML) method that integrates spectrotemporal in…

Vowel Classification

Learning Mid-Level Auditory Codes from Natural Sound Statistics

2017-10-15

Interaction with the world requires an organism to transform sensory signals into representations in which behaviorally meaningful properties of the environment are made explicit. These representations are derived throug…

Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition

2026-02-03 · Xu Zhang, Longbing Cao, Runze Yang, Zhangkai Wu arxiv

Speech emotion recognition (SER) is essential for humanoid robot tasks such as social robotic interactions and robotic psychological diagnosis, where interpretable and efficient models are critical for safety and perform…

Speech Emotion RecognitionRepresentation Learning

Event-driven Spectrotemporal Feature Extraction and Classification using a Silicon Cochlea Model

2022-12-14 · Ying Xu, Samalika Perera, Yeshwanth Bethi, Saeed Afshar 외

This paper presents a reconfigurable digital implementation of an event-based binaural cochlear system on a Field Programmable Gate Array (FPGA). It consists of a pair of the Cascade of Asymmetric Resonators with Fast Ac…