paper-with-me

홈 › Papers

SE-AGCNet: An End-to-End Framework for Joint Speech Enhancement and Loudness Control in Meeting Scenarios

2026-06-24 · Jinming Zhang, Wei Rao, Xionghu Zhong, Eng Siong Chng arxiv

Conventional audio pipelines typically treat speech enhancement (SE) and automatic gain control (AGC) as discrete modules, which often limits overall performance. For instance, applying AGC before SE may inadvertently amplify background noise, while prioritizing SE tends to over-suppress low-volume speech. To address these limitations, we propose SE-AGCNet, an end-to-end framework that jointly optimizes SE and AGC. Tailored for meeting scenarios with significant volume variations, SE-AGCNet leverages the synergy between the two tasks: SE preserves quiet speech, thereby facilitating effective volume adjustment by the AGC component. Furthermore, we propose a specialized data simulation pipeline, SE-AGC-DataGen, and incorporate standardized loudness evaluation metrics: integrated loudness (LUFS), short-term loudness (St LUFS), and LRA. Experiments show that SE-AGCNet consistently achieves target loudness while improving speech quality and ASR accuracy over competitive baselines.

📄 PDF Abstract BibTeX arXiv:2606.25959

Code (0)

등록된 구현이 없습니다.

Tasks

Speech Enhancement

Similar Papers 제목 키워드 기반

Speech Loudness in Broadcasting and Streaming

2024-05-27 · Matteo Torcoli, Mhd Modar Halimeh, Thomas Leitz, Yannik Grewe 외

The introduction and regulation of loudness in broadcasting and streaming brought clear benefits to the audience, e.g., a level of uniformity across programs and channels. Yet, speech loudness is frequently reported as b…

AnCoGen: Analysis, Control and Generation of Speech with a Masked Autoencoder

2025-01-09 · Samir Sadok, Simon Leglaive, Laurent Girin, Gaël Richard 외

This article introduces AnCoGen, a novel method that leverages a masked autoencoder to unify the analysis, control, and generation of speech signals within a single model. AnCoGen can analyze speech by estimating key att…

Pitch ClassificationPitch controlResynthesisSpeech Enhancement+1

SEANet: A Multi-modal Speech Enhancement Network

2020-09-04 · Marco Tagliasacchi, Yunpeng Li, Karolis Misiunas, Dominik Roblek

We explore the possibility of leveraging accelerometer data to perform speech enhancement in very noisy conditions. Although it is possible to only partially reconstruct user's speech from the accelerometer, the latter p…

Speech Enhancement

Universal Speech Enhancement with Score-based Diffusion

2022-06-07 · Joan Serrà, Santiago Pascual, Jordi Pons, R. Oguz Araz 외

Removing background noise from speech audio has been the subject of considerable effort, especially in recent years due to the rise of virtual communication and amateur recordings. Yet background noise is not the only un…

Speech Enhancement

Fighting Game Commentator with Pitch and Loudness Adjustment Utilizing Highlight Cues

2021-08-18 · Junjie H. Xu, Zhou Fang, Qihang Chen, Satoru Ohno 외

This paper presents a commentator for providing real-time game commentary in a fighting game. The commentary takes into account highlight cues, obtained by analyzing scenes during gameplay, as input to adjust the pitch a…

text-to-speechText to Speech