paper-with-me

홈 › Papers

Real-time Audio Video Enhancement \\with a Microphone Array and Headphones

2023-03-02 · Jacob Kealey, Anthony Gosselin, Étienne Deshaies-Samson, Francis Cardinal, Félix Ducharme-Turcotte, Olivier Bergeron, Amélie Rioux-Joyal, Jérémy Bélec, François Grondin

This paper presents a complete hardware and software pipeline for real-time speech enhancement in noisy and reverberant conditions. The device consists of a microphone array and a camera mounted on eyeglasses, connected to an embedded system that enhances speech and plays back the audio in headphones, with a latency of maximum 120 msec. The proposed approach relies on face detection, tracking and verification to enhance the speech of a target speaker using a beamformer and a postfiltering neural network. Results demonstrate the feasibility of the approach, and opens the door to the exploration and validation of a wide range of beamformer and speech enhancement methods for real-time speech enhancement.

📄 PDF Abstract BibTeX arXiv:2303.00949

Code (0)

등록된 구현이 없습니다.

Tasks

Face DetectionSpeech EnhancementVideo Enhancement

Similar Papers 제목 키워드 기반

Cooperative Audio Source Separation and Enhancement Using Distributed Microphone Arrays and Wearable Devices

2019-12-10

Augmented listening devices such as hearing aids often perform poorly in noisy and reverberant environments with many competing sound sources. Large distributed microphone arrays can improve performance, but data from re…

Audio Source Separation

Real-Time System for Audio-Visual Target Speech Enhancement

2025-09-25 · T. Aleksandra Ma, Sile Yin, Li-Chia Yang, Shuo Zhang arxiv

We present a live demonstration for RAVEN, a real-time audio-visual speech enhancement system designed to run entirely on a CPU. In single-channel, audio-only settings, speech enhancement is traditionally approached as t…

Audio-Visual Speech RecognitionSpeech Enhancement

Utterance-Wise Meeting Transcription System Using Asynchronous Distributed Microphones

2020-07-31 · Shota Horiguchi, Yusuke Fujita, Kenji Nagamatsu

A novel framework for meeting transcription using asynchronous microphones is proposed in this paper. It consists of audio synchronization, speaker diarization, utterance-wise speech enhancement using guided source separ…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speaker-diarizationSpeaker Diarization+3

A Causal U-net based Neural Beamforming Network for Real-Time Multi-Channel Speech Enhancement

2021-08-01 · INTERSPEECH 2021 2021 8 · Xinlei Ren, Xu Zhang, LianWu Chen, Xiguang Zheng 외

People are meeting through video conferencing more often. While single channel speech enhancement techniques are useful for the individual participants, the speech quality will be significantly degraded in large meeting …

CPUSpeech Enhancement

Multi-Microphone and Multi-Modal Emotion Recognition in Reverberant Environment

2024-09-14 · Ohad Cohen, Gershon Hazan, Sharon Gannot

This paper presents a Multi-modal Emotion Recognition (MER) system designed to enhance emotion recognition accuracy in challenging acoustic conditions. Our approach combines a modified and extended Hierarchical Token-sem…

Emotion Recognition