Speech Enhancement for Virtual Meetings on Cellular Networks
We study speech enhancement using deep learning (DL) for virtual meetings on cellular devices, where transmitted speech has background noise and transmission loss that affects speech quality. Since the Deep Noise Suppression (DNS) Challenge dataset does not contain practical disturbance, we collect a transmitted DNS (t-DNS) dataset using Zoom Meetings over T-Mobile network. We select two baseline models: Demucs and FullSubNet. The Demucs is an end-to-end model that takes time-domain inputs and outputs time-domain denoised speech, and the FullSubNet takes time-frequency-domain inputs and outputs the energy ratio of the target speech in the inputs. The goal of this project is to enhance the speech transmitted over the cellular networks using deep learning models.
Code (1)
Tasks
Deep LearningSpeech EnhancementSimilar Papers 제목 키워드 기반
Improving Meeting Inclusiveness using Speech Interruption Analysis
Meetings are a pervasive method of communication within all types of companies and organizations, and using remote collaboration systems to conduct meetings has increased dramatically since the COVID-19 pandemic. However…
SE Territory: Monaural Speech Enhancement Meets the Fixed Virtual Perceptual Space Mapping
Monaural speech enhancement has achieved remarkable progress recently. However, its performance has been constrained by the limited spatial cues available at a single microphone. To overcome this limitation, we introduce…
Multi-Task LearningSpeech EnhancementA Speech Recognizer for Frisian/Dutch Council Meetings
We developed a bilingual Frisian/Dutch speech recognizer for council meetings in Fryslân (the Netherlands). During these meetings both Frisian and Dutch are spoken, and code switching between both languages shows up freq…
Deepfake in the Metaverse: Security Implications for Virtual Gaming, Meetings, and Offices
The metaverse has gained significant attention from various industries due to its potential to create a fully immersive and interactive virtual world. However, the integration of deepfakes in the metaverse brings serious…
Face SwappingMMS-MSG: A Multi-purpose Multi-Speaker Mixture Signal Generator
The scope of speech enhancement has changed from a monolithic view of single, independent tasks, to a joint processing of complex conversational speech recordings. Training and evaluation of these single tasks requires s…
Speech Enhancement