paper-with-me

Papers

Multimodal Speech Enhancement Using Burst Propagation

2022-09-07 · Mohsin Raza, Leandro A. Passos, Ahmed Khubaib, Ahsan Adeel

This paper proposes the MBURST, a novel multimodal solution for audio-visual speech enhancements that consider the most recent neurological discoveries regarding pyramidal cells of the prefrontal cortex and other brain regions. The so-called burst propagation implements several criteria to address the credit assignment problem in a more biologically plausible manner: steering the sign and magnitude of plasticity through feedback, multiplexing the feedback and feedforward information across layers through different weight connections, approximating feedback and feedforward connections, and linearizing the feedback signals. MBURST benefits from such capabilities to learn correlations between the noisy signal and the visual stimuli, thus attributing meaning to the speech by amplifying relevant information and suppressing noise. Experiments conducted over a Grid Corpus and CHiME3-based dataset show that MBURST can reproduce similar mask reconstructions to the multimodal backpropagation-based baseline while demonstrating outstanding energy efficiency management, reducing the neuron firing rates to values up to \textbf{$70\%$} lower. Such a feature implies more sustainable implementations, suitable and desirable for hearing aids or any other similar embedded systems.

📄 PDF Abstract BibTeX arXiv:2209.03275

Code (0)

등록된 구현이 없습니다.

Tasks

ManagementSpeech Enhancement

Similar Papers 제목 키워드 기반

Multimodal Spatiotemporal-Frequency Fusion with Peak Enhancement for Cellular Traffic Forecasting

2026-07-08 · Qingzhong Li, Yue Hu, Hui Ma, Yajun Zhang 외 arxiv

Accurate forecasting of cellular network traffic is essential for network planning, resource allocation, and quality-of-service assurance in modern mobile communication systems. Real-world traffic often exhibits bursty e…

Time-Frequency Weighted Losses for Phoneme Reconstruction in DNN-Based Speech Enhancement

2026-06-19 · Nasser-Eddine Monir, Paul Magron, Romain Serizel arxiv

Conventional training losses for speech enhancement based on the signal-to-distortion ratio (SDR) treat all time-frequency (TF) regions uniformly, overlooking the fine-grained spectral cues that are relevant to specific …

Speech Enhancement

Bone-conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models

2026-01-18 · Sina Khanagha, Bunlong Lay, Timo Gerkmann arxiv

Single-channel speech enhancement models face significant performance degradation in extremely noisy environments. While prior work has shown that complementary bone-conducted speech can guide enhancement, effective inte…

Speech Enhancement

Burstormer: Burst Image Restoration and Enhancement Transformer

2023-04-03 · CVPR 2023 1 · Akshay Dudhane, Syed Waqas Zamir, Salman Khan, Fahad Shahbaz Khan 외

On a shutter press, modern handheld cameras capture multiple images in rapid succession and merge them to generate a single image. However, individual frames in a burst are misaligned due to inevitable motions and contai…

DenoisingImage RestorationSuper-Resolution

Audio-Visual Speech Enhancement Using Multimodal Deep Convolutional Neural Networks

2017-03-30 · Jen-Cheng Hou, Syu-Siang Wang, Ying-Hui Lai, Yu Tsao 외

Speech enhancement (SE) aims to reduce noise in speech signals. Most SE techniques focus only on addressing audio information. In this work, inspired by multimodal learning, which utilizes data from different modalities,…

DecoderMulti-Task LearningSpeech Enhancement