paper-with-me

홈 › Papers

SEANet: A Multi-modal Speech Enhancement Network

2020-09-04 · Marco Tagliasacchi, Yunpeng Li, Karolis Misiunas, Dominik Roblek

We explore the possibility of leveraging accelerometer data to perform speech enhancement in very noisy conditions. Although it is possible to only partially reconstruct user's speech from the accelerometer, the latter provides a strong conditioning signal that is not influenced from noise sources in the environment. Based on this observation, we feed a multi-modal input to SEANet (Sound EnhAncement Network), a wave-to-wave fully convolutional model, which adopts a combination of feature losses and adversarial losses to reconstruct an enhanced version of user's speech. We trained our model with data collected by sensors mounted on an earbud and synthetically corrupted by adding different kinds of noise sources to the audio signal. Our experimental results demonstrate that it is possible to achieve very high quality results, even in the case of interfering speech at the same level of loudness. A sample of the output produced by our model is available at https://google-research.github.io/seanet/multimodal/speech.

📄 PDF Abstract BibTeX arXiv:2009.02095

Code (1)

google-research/seanet 공식 구현

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Real-time Speech Frequency Bandwidth Extension

2020-10-21

In this paper we propose a lightweight model for frequency bandwidth extension of speech signals, increasing the sampling frequency from 8kHz to 16kHz while restoring the high frequency content to a level almost indistin…

Bandwidth ExtensionCPU

Audio-Visual Target Speaker Extraction with Reverse Selective Auditory Attention

2024-04-29 · Ruijie Tao, Xinyuan Qian, Yidi Jiang, Junjie Li 외

Audio-visual target speaker extraction (AV-TSE) aims to extract the specific person's speech from the audio mixture given auxiliary visual cues. Previous methods usually search for the target voice through speech-lip syn…

Target Speaker Extraction

Lightweight Salient Object Detection in Optical Remote-Sensing Images via Semantic Matching and Edge Alignment

2023-01-07 · Gongyang Li, Zhi Liu, Xinpeng Zhang, Weisi Lin

Recently, relying on convolutional neural networks (CNNs), many methods for salient object detection in optical remote sensing images (ORSI-SOD) are proposed. However, most methods ignore the huge parameters and computat…

Decoderobject-detectionObject DetectionSalient Object Detection

AudioPaLM: A Large Language Model That Can Speak and Listen

2023-06-22 · Paul K. Rubenstein, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna 외

We introduce AudioPaLM, a large language model for speech understanding and generation. AudioPaLM fuses text-based and speech-based language models, PaLM-2 [Anil et al., 2023] and AudioLM [Borsos et al., 2022], into a un…

Language ModelingLanguage ModellingLarge Language Modelspeech-recognition+5

Bone-conduction Guided Multimodal Speech Enhancement with Conditional Diffusion Models

2026-01-18 · Sina Khanagha, Bunlong Lay, Timo Gerkmann arxiv

Single-channel speech enhancement models face significant performance degradation in extremely noisy environments. While prior work has shown that complementary bone-conducted speech can guide enhancement, effective inte…

Speech Enhancement