paper-with-me

홈 › Papers

WaveLLDM: Design and Development of a Lightweight Latent Diffusion Model for Speech Enhancement and Restoration

2025-08-28 · Kevin Putra Santoso, Rizka Wakhidatus Sholikah, Raden Venantius Hari Ginardi arxiv

High-quality audio is essential in a wide range of applications, including online communication, virtual assistants, and the multimedia industry. However, degradation caused by noise, compression, and transmission artifacts remains a major challenge. While diffusion models have proven effective for audio restoration, they typically require significant computational resources and struggle to handle longer missing segments. This study introduces WaveLLDM (Wave Lightweight Latent Diffusion Model), an architecture that integrates an efficient neural audio codec with latent diffusion for audio restoration and denoising. Unlike conventional approaches that operate in the time or spectral domain, WaveLLDM processes audio in a compressed latent space, reducing computational complexity while preserving reconstruction quality. Empirical evaluations on the Voicebank+DEMAND test set demonstrate that WaveLLDM achieves accurate spectral reconstruction with low Log-Spectral Distance (LSD) scores (0.48 to 0.60) and good adaptability to unseen data. However, it still underperforms compared to state-of-the-art methods in terms of perceptual quality and speech clarity, with WB-PESQ scores ranging from 1.62 to 1.71 and STOI scores between 0.76 and 0.78. These limitations are attributed to suboptimal architectural tuning, the absence of fine-tuning, and insufficient training duration. Nevertheless, the flexible architecture that combines a neural audio codec and latent diffusion model provides a strong foundation for future development.

📄 PDF Abstract BibTeX arXiv:2508.21153

Code (0)

등록된 구현이 없습니다.

Tasks

Spectral ReconstructionSpeech Enhancement

Similar Papers 제목 키워드 기반

PeLAP-A: Adaptive Latent Pruning for Lightweight Latent Diffusion Models

2026-06-22 · Kissa Zahra, Zaib Un Nisa arxiv

Latent diffusion models achieve strong generative performance by operating in a compressed latent space produced by a variational autoencoder (VAE). However, it remains unclear whether all latent channels contribute equa…

Diffusion Timbre Transfer Via Mutual Information Guided Inpainting

2026-01-03 · Ching Ho Lee, Javier Nistal, Stefan Lattner, Marco Pasini 외 arxiv

We study timbre transfer as an inference-time editing problem for music audio. Starting from a strong pre-trained latent diffusion model, we introduce a lightweight procedure that requires no additional training: (i) a d…

Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition

2024-03-21 · Sihyun Yu, Weili Nie, De-An Huang, Boyi Li 외

Video diffusion models have recently made great progress in generation quality, but are still limited by the high memory and computational requirements. This is because current video diffusion models often attempt to pro…

Video Generation

NanoSD: Edge Efficient Foundation Model for Real Time Image Restoration

2026-01-14 · Subhajit Sanyal, Srinivas Soumitri Miriyala, Akshay Janardan Bankar, Manjunath Arveti 외 arxiv

Latent diffusion models such as Stable Diffusion 1.5 offer strong generative priors that are highly valuable for image restoration, yet their full pipelines remain too computationally heavy for deployment on edge devices…

Monocular Depth EstimationImage Super-ResolutionImage RestorationImage Deblurring

LatentHDR: Decoupling Exposure from Diffusion via Conditional Latent-to-Latent Mapping for Text/Image-to-Panoramic HDR

2026-05-11 · Pedram Fekri, WenChen Li, William Chen, Peter Altamirano arxiv

High Dynamic Range (HDR) generation remains challenging for generative models, which are largely limited to low dynamic range outputs. Recent diffusionbased approaches approximate HDR by generating multiple exposure-cond…

Scene Generation