paper-with-me

Papers

Multi-encoder attention-based architectures for sound recognition with partial visual assistance

2022-09-26 · Wim Boes, Hugo Van hamme

Large-scale sound recognition data sets typically consist of acoustic recordings obtained from multimedia libraries. As a consequence, modalities other than audio can often be exploited to improve the outputs of models designed for associated tasks. Frequently, however, not all contents are available for all samples of such a collection: For example, the original material may have been removed from the source platform at some point, and therefore, non-auditory features can no longer be acquired. We demonstrate that a multi-encoder framework can be employed to deal with this issue by applying this method to attention-based deep learning systems, which are currently part of the state of the art in the domain of sound recognition. More specifically, we show that the proposed model extension can successfully be utilized to incorporate partially available visual information into the operational procedures of such networks, which normally only use auditory features during training and inference. Experimentally, we verify that the considered approach leads to improved predictions in a number of evaluation scenarios pertaining to audio tagging and sound event detection. Additionally, we scrutinize some properties and limitations of the presented technique.

📄 PDF Abstract BibTeX arXiv:2209.12826

Code (0)

등록된 구현이 없습니다.

Tasks

Audio TaggingEvent DetectionSound Event Detection

Similar Papers 제목 키워드 기반

Exploiting Attention-based Sequence-to-Sequence Architectures for Sound Event Localization

2021-02-28 · Christopher Schymura, Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita 외

Sound event localization frameworks based on deep neural networks have shown increased robustness with respect to reverberation and noise in comparison to classical parametric approaches. In particular, recurrent archite…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition

Synthesizer Preset Interpolation using Transformer Auto-Encoders

2022-10-27 · Gwendal Le Vaillant, Thierry Dutoit

Sound synthesizers are widespread in modern music production but they increasingly require expert skills to be mastered. This work focuses on interpolation between presets, i.e., sets of values of all sound synthesis par…

Deblurring Masked Autoencoder is Better Recipe for Ultrasound Image Recognition

2023-06-14 · Qingbo Kang, Jun Gao, Kang Li, Qicheng Lao

Masked autoencoder (MAE) has attracted unprecedented attention and achieves remarkable performance in many vision tasks. It reconstructs random masked image patches (known as proxy task) during pretraining and learns mea…

Deblurringimage-classificationImage Classification

On the Relevance of Temporal Features for Medical Ultrasound Video Recognition

2023-10-16 · D. Hudson Smith, John Paul Lineberger, George H. Baker

Many medical ultrasound video recognition tasks involve identifying key anatomical features regardless of when they appear in the video suggesting that modeling such tasks may not benefit from temporal features. Correspo…

Video Recognition

Multi-encoder multi-resolution framework for end-to-end speech recognition

2018-11-12 · Ruizhi Li, Xiaofei Wang, Sri Harish Mallidi, Takaaki Hori 외

Attention-based methods and Connectionist Temporal Classification (CTC) network have been promising research directions for end-to-end Automatic Speech Recognition (ASR). The joint CTC/Attention model has achieved great …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech Recognition