paper-with-me

홈 › Papers

Biomimetic Frontend for Differentiable Audio Processing

2024-09-13 · Ruolan Leslie Famularo, Dmitry N. Zotkin, Shihab A. Shamma, Ramani Duraiswami

While models in audio and speech processing are becoming deeper and more end-to-end, they as a consequence need expensive training on large data, and are often brittle. We build on a classical model of human hearing and make it differentiable, so that we can combine traditional explainable biomimetic signal processing approaches with deep-learning frameworks. This allows us to arrive at an expressive and explainable model that is easily trained on modest amounts of data. We apply this model to audio processing tasks, including classification and enhancement. Results show that our differentiable model surpasses black-box approaches in terms of computational efficiency and robustness, even with little training data. We also discuss other potential applications.

📄 PDF Abstract BibTeX arXiv:2409.08997

Code (1)

pirl-lab/diffAudNeuro 공식 구현 jax

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

EfficientLEAF: A Faster LEarnable Audio Frontend of Questionable Use

2022-07-12 · Jan Schlüter, Gerald Gutenbrunner

In audio classification, differentiable auditory filterbanks with few parameters cover the middle ground between hard-coded spectrograms and raw audio. LEAF (arXiv:2101.08596), a Gabor-based filterbank combined with Per-…

Audio ClassificationClassificationInstrument RecognitionPitch Classification+1

Online Audio-Visual Autoregressive Speaker Extraction

2025-06-02 · Zexu Pan, Wupeng Wang, Shengkui Zhao, Chong Zhang 외

This paper proposes a novel online audio-visual speaker extraction model. In the streaming regime, most studies optimize the audio network only, leaving the visual frontend less explored. We first propose a lightweight v…

Learnable Frontends that do not Learn: Quantifying Sensitivity to Filterbank Initialisation

2023-02-20 · Mark Anderson, Tomi Kinnunen, Naomi Harte

While much of modern speech and audio processing relies on deep neural networks trained using fixed audio representations, recent studies suggest great potential in acoustic frontends learnt jointly with a backend. In th…

Action DetectionActivity DetectionSensitivity

A Universal Learnable Audio Frontend

2021-01-01 · ICLR 2021 1 · Neil Zeghidour, Olivier Teboul, Félix de Chaumont Quitry, Marco Tagliasacchi

Mel-filterbanks are fixed, engineered audio features which emulate human perception and have lived through the history of audio understanding up to today. However, their undeniable qualities are counterbalanced by the fu…

Audio Classification

LEAF: A Learnable Frontend for Audio Classification

2021-01-21 · Neil Zeghidour, Olivier Teboul, Félix de Chaumont Quitry, Marco Tagliasacchi

Mel-filterbanks are fixed, engineered audio features which emulate human perception and have been used through the history of audio understanding up to today. However, their undeniable qualities are counterbalanced by th…

Audio ClassificationClassificationGeneral Classification