paper-with-me

홈 › Papers

Speech Loudness in Broadcasting and Streaming

2024-05-27 · Matteo Torcoli, Mhd Modar Halimeh, Thomas Leitz, Yannik Grewe, Michael Kratschmer, Bernhard Neugebauer, Adrian Murtaza, Harald Fuchs, Emanuël A. P. Habets

The introduction and regulation of loudness in broadcasting and streaming brought clear benefits to the audience, e.g., a level of uniformity across programs and channels. Yet, speech loudness is frequently reported as being too low in certain passages, which can hinder the full understanding and enjoyment of movies and TV programs. This paper proposes expanding the set of loudness-based measures typically used in the industry. We focus on speech loudness, and we show that, when clean speech is not available, Deep Neural Networks (DNNs) can be used to isolate the speech signal and so to accurately estimate speech loudness, providing a more precise estimate compared to speech-gated loudness. Moreover, we define critical passages, i.e., passages in which speech is likely to be hard to understand. Critical passages are defined based on the local Speech Loudness Deviation (SLD) and the local Speech-to-Background Loudness Difference (SBLD), as SLD and SBLD significantly contribute to intelligibility and listening effort. In contrast to other more comprehensive measures of intelligibility and listening effort, SLD and SBLD can be straightforwardly measured, are intuitive, and, most importantly, can be easily controlled by adjusting the speech level in the mix or by enabling personalization at the user's end. Finally, examples are provided that show how the detection of critical passages can support the evaluation and control of the speech signal during and after content production.

📄 PDF Abstract BibTeX arXiv:2405.17364

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

SE-AGCNet: An End-to-End Framework for Joint Speech Enhancement and Loudness Control in Meeting Scenarios

2026-06-24 · Jinming Zhang, Wei Rao, Xionghu Zhong, Eng Siong Chng arxiv

Conventional audio pipelines typically treat speech enhancement (SE) and automatic gain control (AGC) as discrete modules, which often limits overall performance. For instance, applying AGC before SE may inadvertently am…

Speech Enhancement

MPEG-H Audio for Improving Accessibility in Broadcasting and Streaming

2019-09-25

Broadcasting and streaming services still suffer from various levels of accessibility barriers for a significant portion of the population, limiting the access to information and culture, and in the most severe cases lim…

Cultural Vocal Bursts Intensity Prediction

Fighting Game Commentator with Pitch and Loudness Adjustment Utilizing Highlight Cues

2021-08-18 · Junjie H. Xu, Zhou Fang, Qihang Chen, Satoru Ohno 외

This paper presents a commentator for providing real-time game commentary in a fighting game. The commentary takes into account highlight cues, obtained by analyzing scenes during gameplay, as input to adjust the pitch a…

text-to-speechText to Speech

Prosody Transfer in Neural Text to Speech Using Global Pitch and Loudness Features

2019-11-21 · Siddharth Gururani, Kilol Gupta, Dhaval Shah, Zahra Shakeri 외

This paper presents a simple yet effective method to achieve prosody transfer from a reference speech signal to synthesized speech. The main idea is to incorporate well-known acoustic correlates of prosody such as pitch …

text-to-speechText to Speech

AI-Generated Music Detection in Broadcast Monitoring

2026-02-06 · David López-Ayala, Asier Cabello, Pablo Zinemanas, Emilio Molina 외 arxiv

AI music generators have advanced to the point where their outputs are often indistinguishable from human compositions. While detection methods have emerged, they are typically designed and validated in music streaming c…