paper-with-me

홈 › Papers

Comparison of Time-Frequency Representations for Environmental Sound Classification using Convolutional Neural Networks

2017-06-22 · M. Huzaifah

Recent successful applications of convolutional neural networks (CNNs) to audio classification and speech recognition have motivated the search for better input representations for more efficient training. Visual displays of an audio signal, through various time-frequency representations such as spectrograms offer a rich representation of the temporal and spectral structure of the original signal. In this letter, we compare various popular signal processing methods to obtain this representation, such as short-time Fourier transform (STFT) with linear and Mel scales, constant-Q transform (CQT) and continuous Wavelet transform (CWT), and assess their impact on the classification performance of two environmental sound datasets using CNNs. This study supports the hypothesis that time-frequency representations are valuable in learning useful features for sound classification. Moreover, the actual transformation used is shown to impact the classification accuracy, with Mel-scaled STFT outperforming the other discussed methods slightly and baseline MFCC features to a large degree. Additionally, we observe that the optimal window size during transformation is dependent on the characteristics of the audio signal and architecturally, 2D convolution yielded better results in most cases compared to 1D.

📄 PDF Abstract BibTeX arXiv:1706.07156

Code (0)

등록된 구현이 없습니다.

Tasks

Audio ClassificationClassificationEnvironmental Sound ClassificationGeneral ClassificationSound Classificationspeech-recognitionSpeech Recognition

Methods 이 논문이 사용한 방법론

Convolution A convolution is a type of matrix operation, consisting of a kernel, a small matrix of weights, that slides over input data performing element-wise multiplication with the…

Similar Papers 제목 키워드 기반

Diversity-Robust Acoustic Feature Signatures Based on Multiscale Fractal Dimension for Similarity Search of Environmental Sounds

2021-02-05 · Motohiro Sunouchi, Masaharu Yoshioka

This paper proposes new acoustic feature signatures based on the multiscale fractal dimension (MFD), which are robust against the diversity of environmental sounds, for the content-based similarity search. The diversity …

Density EstimationDiversity

Environmental Sound Classification with Parallel Temporal-spectral Attention

2019-12-14 · Helin Wang, Yuexian Zou, Dading Chong, Wenwu Wang

Convolutional neural networks (CNN) are one of the best-performing neural network architectures for environmental sound classification (ESC). Recently, temporal attention mechanisms have been used in CNN to capture the u…

Acoustic Scene ClassificationAudio ClassificationClassificationEnvironmental Sound Classification+3

BEAT2AASIST model with layer fusion for ESDD 2026 Challenge

2025-12-17 · Sanghyeok Chung, Eujin Kim, Donggun Kim, Gaeun Heo 외 arxiv

Recent advances in audio generation have increased the risk of realistic environmental sound manipulation, motivating the ESDD 2026 Challenge as the first large-scale benchmark for Environmental Sound Deepfake Detection …

DeepFake DetectionData AugmentationAudio Generation

Environmental Sound Extraction Using Onomatopoeic Words

2021-12-01 · Yuki Okamoto, Shota Horiguchi, Masaaki Yamamoto, Keisuke Imoto 외

An onomatopoeic word, which is a character sequence that phonetically imitates a sound, is effective in expressing characteristics of sound such as duration, pitch, and timbre. We propose an environmental-sound-extractio…

From Sound Representation to Model Robustness

2020-07-27 · Mohammad Esmaeilpour, Patrick Cardinal, Alessandro Lameiras Koerich

In this paper, we investigate the impact of different standard environmental sound representations (spectrograms) on the recognition performance and adversarial attack robustness of a victim residual convolutional neural…

Adversarial AttackAdversarial RobustnessBenchmarkingmodel