Papers Audio Signal Processing
“Audio Signal Processing” 태그가 달린 논문 70편 · 필터 해제
Full Attention Bidirectional Deep Learning Structure for Single Channel Speech Enhancement
As the cornerstone of other important technologies, such as speech recognition and speech synthesis, speech enhancement is a critical area in audio signal processing. In this paper, a new deep learning structure for spee…
Audio Signal ProcessingSpeech Enhancementspeech-recognitionSpeech Recognition+1Blind Identification of State-Space Models in Physical Coordinates
Blind identification is popular for modeling a system without the input information, such as in the research areas of structural health monitoring and audio signal processing. Existing blind identification methods have b…
Audio Signal ProcessingState Space ModelsStructural Health MonitoringDifferentiable Signal Processing With Black-Box Audio Effects
We present a data-driven approach to automate audio signal processing by incorporating stateful third-party, audio effects as layers within a deep neural network. We then train a deep encoder to analyze input audio and c…
Audio Signal ProcessingVisualization of Linear Operations in the Spherical Harmonics Domain
Linear operations on coefficients in the spherical harmonics (SH) transform domain that again yield SH-domain coefficients are an important toolset in many disciplines of research and engineering. They comprise rotations…
Audio Signal ProcessingBeamLearning: an end-to-end Deep Learning approach for the angular localization of sound sources using raw multichannel acoustic pressure data
Sound sources localization using multichannel signal processing has been a subject of active research for decades. In recent years, the use of deep learning in audio signal processing has allowed to drastically improve p…
Audio Signal ProcessingBIG-bench Machine LearningComputational EfficiencyGPUDeepSpectrumLite: A Power-Efficient Transfer Learning Framework for Embedded Speech and Audio Processing from Decentralised Data
Deep neural speech and audio processing systems have a large number of trainable parameters, a relatively complex architecture, and require a vast amount of training data and computational power. These constraints make i…
Audio Signal ProcessingTransfer LearningL3DAS21 Challenge: Machine Learning for 3D Audio Signal Processing
The L3DAS21 Challenge is aimed at encouraging and fostering collaborative research on machine learning for 3D audio signal processing, with particular focus on 3D speech enhancement (SE) and 3D sound localization and det…
Audio Signal ProcessingBIG-bench Machine LearningSpeech EnhancementMelon Playlist Dataset: a public dataset for audio-based playlist generation and music tagging
One of the main limitations in the field of audio signal processing is the lack of large public datasets with audio representations and high-quality annotations due to restrictions of copyrighted commercial music. We pre…
Audio Signal ProcessingCollaborative FilteringInformation RetrievalMetric Learning+4A Survey on Deep Reinforcement Learning for Audio-Based Applications
Deep reinforcement learning (DRL) is poised to revolutionise the field of artificial intelligence (AI) by endowing autonomous systems with high levels of understanding of the real world. Currently, deep learning (DL) is …
Audio Signal ProcessingDeep Reinforcement Learningreinforcement-learningReinforcement Learning+2Upsampling artifacts in neural audio synthesis
A number of recent advances in neural audio synthesis rely on upsampling layers, which can introduce undesired artifacts. In computer vision, upsampling artifacts have been studied and are known as checkerboard artifacts…
Audio Signal ProcessingAudio SynthesisContextualized Attention-based Knowledge Transfer for Spoken Conversational Question Answering
Spoken conversational question answering (SCQA) requires machines to model complex dialogue flow given the speech utterances and text corpora. Different from traditional text question answering (QA) tasks, SCQA involves …
Audio Signal ProcessingConversational Question AnsweringKnowledge DistillationQuestion Answering+1Utterance Clustering Using Stereo Audio Channels
Utterance clustering is one of the actively researched topics in audio signal processing and machine learning. This study aims to improve the performance of utterance clustering by processing multichannel (stereo) audio …
Audio Signal ProcessingClusteringSpeaker Diarizationwav2shape: Hearing the Shape of a Drum Machine
Disentangling and recovering physical attributes, such as shape and material, from a few waveform examples is a challenging inverse problem in audio signal processing, with numerous applications in musical acoustics as w…
Audio Signal ProcessingPrivate Speech Classification with Secure Multiparty Computation
Deep learning in audio signal processing, such as human voice audio signal classification, is a rich application area of machine learning. Legitimate use cases include voice authentication, gunfire detection, and emotion…
Audio ClassificationAudio Signal ProcessingClassificationEmotion Recognition+2Exploring Quality and Generalizability in Parameterized Neural Audio Effects
Deep neural networks have shown promise for music audio signal processing applications, often surpassing prior approaches, particularly as end-to-end models in the waveform domain. Yet results to date have tended to be c…
Audio Signal ProcessingComputational EfficiencySignalTrain: Profiling Audio Compressors with Deep Neural Networks
In this work we present a data-driven approach for predicting the behavior of (i.e., profiling) a given non-linear audio signal processing effect (henceforth "audio effect"). Our objective is to learn a mapping function …
Audio Effects ModelingAudio Signal ProcessingDeep Learning for Audio Signal Processing
Given the recent surge in developments of deep learning, this article provides a review of the state-of-the-art deep learning techniques for audio signal processing. Speech, music, and environmental sound processing are …
Audio Signal ProcessingAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Deep Learning+5Practical Hidden Voice Attacks against Speech and Speaker Recognition Systems
Voice Processing Systems (VPSes), now widely deployed, have been made significantly more accurate through the application of recent advances in machine learning. However, adversarial machine learning has similarly advanc…
Audio Signal ProcessingBIG-bench Machine LearningSpeaker RecognitionLow-rankness of Complex-valued Spectrogram and Its Application to Phase-aware Audio Processing
Low-rankness of amplitude spectrograms has been effectively utilized in audio signal processing methods including non-negative matrix factorization. However, such methods have a fundamental limitation owing to their ampl…
Audio DenoisingAudio Signal ProcessingDenoisingEnd-to-End Probabilistic Inference for Nonstationary Audio Analysis
A typical audio signal processing pipeline includes multiple disjoint analysis stages, including calculation of a time-frequency representation followed by spectrogram-based feature analysis. We show how time-frequency a…
Audio Signal Processingregression