paper-with-me

Papers

Complementing Handcrafted Features with Raw Waveform Using a Light-weight Auxiliary Model

2021-09-06 · Zhongwei Teng, Quchen Fu, Jules White, Maria Powell, Douglas C. Schmidt

An emerging trend in audio processing is capturing low-level speech representations from raw waveforms. These representations have shown promising results on a variety of tasks, such as speech recognition and speech separation. Compared to handcrafted features, learning speech features via backpropagation provides the model greater flexibility in how it represents data for different tasks theoretically. However, results from empirical study shows that, in some tasks, such as voice spoof detection, handcrafted features are more competitive than learned features. Instead of evaluating handcrafted features and raw waveforms independently, this paper proposes an Auxiliary Rawnet model to complement handcrafted features with features learned from raw waveforms. A key benefit of the approach is that it can improve accuracy at a relatively low computational cost. The proposed Auxiliary Rawnet model is tested using the ASVspoof 2019 dataset and the results from this dataset indicate that a light-weight waveform encoder can potentially boost the performance of handcrafted-features-based encoders in exchange for a small amount of additional computational work.

📄 PDF Abstract BibTeX arXiv:2109.02773

Code (1)

magnumresearchgroup/auxiliaryrawnet pytorch

Tasks

speech-recognitionSpeech RecognitionSpeech Separation

Similar Papers 제목 키워드 기반

Zero-Shot Parkinson's Disease Detection from Speech: Comparing Large Audio and Language Models

2026-05-24 · Muhammad Ashad Kabir, Sirajam Munira arxiv

Large audio and language models have recently demonstrated zero-shot reasoning capabilities across various domains. However, it remains unclear how the form of audio input, whether handcrafted acoustic features extracted…

End-to-end Audio Deepfake Detection from RAW Waveforms: a RawNet-Based Approach with Cross-Dataset Evaluation

2025-04-29 · Andrea Di Pierno, Luca Guarnera, Dario Allegra, Sebastiano Battiato

Audio deepfakes represent a growing threat to digital security and trust, leveraging advanced generative models to produce synthetic speech that closely mimics real human voices. Detecting such manipulations is especiall…

Audio Deepfake DetectionDeepFake DetectionFace Swapping

Pushing the limits of raw waveform speaker recognition

2022-03-16 · Jee-weon Jung, You Jin Kim, Hee-Soo Heo, Bong-Jin Lee 외

In recent years, speaker recognition systems based on raw waveform inputs have received increasing attention. However, the performance of such systems are typically inferior to the state-of-the-art handcrafted feature-ba…

Self-Supervised LearningSpeaker RecognitionSpeaker Verification

FRAME-C: A knowledge-augmented deep learning pipeline for classifying multi-electrode array electrophysiological signals

2025-05-18 · Nisal Ranasinghe, Dzung Do-Ha, Simon Maksour, Tamasha Malepathirana 외

Amyotrophic lateral sclerosis (ALS) is a fatal neurodegenerative disorder characterized by motor neuron degeneration, with alterations in neural excitability serving as key indicators. Recent advancements in induced plur…

Deep LearningFeature Importance

Multi-Span Acoustic Modelling using Raw Waveform Signals

2019-06-21 · Patrick von Platen, Chao Zhang, Philip Woodland

Traditional automatic speech recognition (ASR) systems often use an acoustic model (AM) built on handcrafted acoustic features, such as log Mel-filter bank (FBANK) values. Recent studies found that AMs with convolutional…

Acoustic ModellingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognition+1