paper-with-me

홈 › Papers

A Hybrid Continuity Loss to Reduce Over-Suppression for Time-domain Target Speaker Extraction

2022-03-31 · Zexu Pan, Meng Ge, Haizhou Li

The speaker extraction algorithm extracts the target speech from a mixture speech containing interference speech and background noise. The extraction process sometimes over-suppresses the extracted target speech, which not only creates artifacts during listening but also harms the performance of downstream automatic speech recognition algorithms. We propose a hybrid continuity loss function for time-domain speaker extraction algorithms to settle the over-suppression problem. On top of the waveform-level loss used for superior signal quality, i.e., SI-SDR, we introduce a multi-resolution delta spectrum loss in the frequency-domain, to ensure the continuity of an extracted speech signal, thus alleviating the over-suppression. We examine the hybrid continuity loss function using a time-domain audio-visual speaker extraction algorithm on the YouTube LRS2-BBC dataset. Experimental results show that the proposed loss function reduces the over-suppression and improves the word error rate of speech recognition on both clean and noisy two-speakers mixtures, without harming the reconstructed speech quality.

📄 PDF Abstract BibTeX arXiv:2203.16843

Code (1)

zexupan/avse_hybrid_loss 공식 구현 pytorch

Tasks

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)speech-recognitionSpeech RecognitionTarget Speaker Extraction

Similar Papers 제목 키워드 기반

Deep Learning for Joint Acoustic Echo and Acoustic Howling Suppression in Hybrid Meetings

2023-05-02 · Hao Zhang, Meng Yu, Dong Yu

Hybrid meetings have become increasingly necessary during the post-COVID period and also brought new challenges for solving audio-related problems. In particular, the interplay between acoustic echo and acoustic howling …

Speech Separation

Artifact Correction for Echo-Planar Imaging at Low-Field and Ultra-Low-Field MRI

2026-05-25 · Sisi Qiao, Yilin Yu, Tiecheng Lin, Yuhao Liu 외 arxiv

Purpose: Echo-planar imaging (EPI) in low-field (LF) and ultra-low-field MRI (ULF) suffers from severe Nyquist ghost artifacts due to odd-even k-space misalignment. This study develops a reference-free artifact correctio…

Rethinking Temporal Object Detection from Robotic Perspectives

2019-12-22 · Xingyu Chen, Zhengxing Wu, Junzhi Yu, Li Wen

Video object detection (VID) has been vigorously studied for years but almost all literature adopts a static accuracy-based evaluation, i.e., average precision (AP). From a robotic perspective, the importance of recall c…

Multi-Object TrackingObjectobject-detectionObject Detection+2

Phase Continuity: Learning Derivatives of Phase Spectrum for Speech Enhancement

2022-02-24 · Doyeon Kim, Hyewon Han, Hyeon-Kyeong Shin, Soo-Whan Chung 외

Modern neural speech enhancement models usually include various forms of phase information in their training loss terms, either explicitly or implicitly. However, these loss terms are typically designed to reduce the dis…

Speech Enhancement

Hybrid Multi-Stage Learning Framework for Edge Detection: A Survey

2025-03-26 · Mark Phil Pacot, Jayno Juventud, Gleen Dalaorao

Edge detection remains a fundamental yet challenging task in computer vision, especially under varying illumination, noise, and complex scene conditions. This paper introduces a Hybrid Multi-Stage Learning Framework that…

Deep LearningEdge DetectionSurvey