paper-with-me

홈 › Papers

DDD: A Perceptually Superior Low-Response-Time DNN-based Declipper

2024-01-08 · Jayeon Yi, Junghyun Koo, Kyogu Lee

Clipping is a common nonlinear distortion that occurs whenever the input or output of an audio system exceeds the supported range. This phenomenon undermines not only the perception of speech quality but also downstream processes utilizing the disrupted signal. Therefore, a real-time-capable, robust, and low-response-time method for speech declipping (SD) is desired. In this work, we introduce DDD (Demucs-Discriminator-Declipper), a real-time-capable speech-declipping deep neural network (DNN) that requires less response time by design. We first observe that a previously untested real-time-capable DNN model, Demucs, exhibits a reasonable declipping performance. Then we utilize adversarial learning objectives to increase the perceptual quality of output speech without additional inference overhead. Subjective evaluations on harshly clipped speech shows that DDD outperforms the baselines by a wide margin in terms of speech quality. We perform detailed waveform and spectral analyses to gain an insight into the output behavior of DDD in comparison to the baselines. Finally, our streaming simulations also show that DDD is capable of sub-decisecond mean response times, outperforming the state-of-the-art DNN approach by a factor of six.

📄 PDF Abstract BibTeX arXiv:2401.03650

Code (1)

stet-stet/ddd 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Image to Image Translation based on Convolutional Neural Network Approach for Speech Declipping

2019-10-26

Clipping, as a current nonlinear distortion, often occurs due to the limited dynamic range of audio recorders. It degrades the speech quality and intelligibility and adversely affects the performances of speech and speak…

Image-to-Image TranslationTranslation

Perceptually Aligning Representations of Music via Noise-Augmented Autoencoders

2025-11-07 · Mathias Rose Bjare, Giorgia Cantisani, Marco Pasini, Stefan Lattner 외 arxiv

We argue that training autoencoders to reconstruct inputs from noised versions of their encodings, when combined with perceptually motivated losses, yields encodings that are structured according to a perceptual hierarch…

APPLADE: Adjustable Plug-and-play Audio Declipper Combining DNN with Sparse Optimization

2022-02-16 · Tomoro Tanaka, Kohei Yatabe, Masahiro Yasuda, Yasuhiro Oikawa

In this paper, we propose an audio declipping method that takes advantages of both sparse optimization and deep learning. Since sparsity-based audio declipping methods have been developed upon constrained optimization, t…

Audio declipping

Mitigating Perceptual Judgment Bias in Multimodal LLM-as-a-Judge via Perceptual Perturbation and Reward Modeling

2026-06-01 · Seojeong Park, Jiho Choi, Junyong Kang, Seonho Lee 외 arxiv

Recent multimodal large language models have demonstrated strong reasoning ability, yet their reliability as automated evaluators remains limited by a critical weakness: when visual evidence conflicts with textual cues, …

Em-Garde: A Propose-Match Framework for Proactive Streaming Video Understanding

2026-03-19 · Yikai Zheng, Xin Ding, Yifan Yang, Shiqi Jiang 외 arxiv

Recent advances in Streaming Video Understanding has enabled a new interaction paradigm where models respond proactively to user queries. Current proactive VideoLLMs rely on per-frame triggering decision making, which su…

Decision Making