paper-with-me

홈 › Papers

Better Together: Dialogue Separation and Voice Activity Detection for Audio Personalization in TV

2023-03-23 · Matteo Torcoli, Emanuël A. P. Habets

In TV services, dialogue level personalization is key to meeting user preferences and needs. When dialogue and background sounds are not separately available from the production stage, Dialogue Separation (DS) can estimate them to enable personalization. DS was shown to provide clear benefits for the end user. Still, the estimated signals are not perfect, and some leakage can be introduced. This is undesired, especially during passages without dialogue. We propose to combine DS and Voice Activity Detection (VAD), both recently proposed for TV audio. When their combination suggests dialogue inactivity, background components leaking in the dialogue estimate are reassigned to the background estimate. A clear improvement of the audio quality is shown for dialogue-free signals, without performance drops when dialogue is active. A post-processed VAD estimate with improved detection accuracy is also generated. It is concluded that DS and VAD can improve each other and are better used together.

📄 PDF Abstract BibTeX arXiv:2303.13453

Code (0)

등록된 구현이 없습니다.

Tasks

Action DetectionActivity Detection

Similar Papers 제목 키워드 기반

Facing the Music: Tackling Singing Voice Separation in Cinematic Audio Source Separation

2024-08-07 · Karn N. Watcharasupat, Chih-Wei Wu, Iroro Orife

Cinematic audio source separation (CASS), as a standalone problem of extracting individual stems from their mixture, is a fairly new subtask of audio source separation. A typical setup of CASS is a three-stem problem, wi…

Audio Source SeparationDecoder

Single-Microphone Speaker Separation and Voice Activity Detection in Noisy and Reverberant Environments

2024-01-07 · Renana Opochinsky, Mordehay Moradi, Sharon Gannot

Speech separation involves extracting an individual speaker's voice from a multi-speaker audio signal. The increasing complexity of real-world environments, where multiple speakers might converse simultaneously, undersco…

Action DetectionActivity DetectionDecoderSpeaker Separation+1

Triadic Multi-party Voice Activity Projection for Turn-taking in Spoken Dialogue Systems

2025-07-10 · Mikey Elmers, Koji Inoue, Divesh Lala, Tatsuya Kawahara arxiv

Turn-taking is a fundamental component of spoken dialogue, however conventional studies mostly involve dyadic settings. This work focuses on applying voice activity projection (VAP) to predict upcoming turn-taking in tri…

Single channel voice separation for unknown number of speakers under reverberant and noisy settings

2020-11-04 · Shlomo E. Chazan, Lior Wolf, Eliya Nachmani, Yossi Adi

We present a unified network for voice separation of an unknown number of speakers. The proposed approach is composed of several separation heads optimized together with a speaker classification branch. The separation is…

ClassificationGeneral Classification

Front-end Diarization for Percussion Separation in Taniavartanam of Carnatic Music Concerts

2021-03-04 · Nauman Dawalatabad, Jilt Sebastian, Jom Kuriakose, C. Chandra Sekhar 외

Instrument separation in an ensemble is a challenging task. In this work, we address the problem of separating the percussive voices in the taniavartanam segments of Carnatic music. In taniavartanam, a number of percussi…