paper-with-me

홈 › Papers

EGGCodec: A Robust Neural Encodec Framework for EGG Reconstruction and F0 Extraction

2025-08-12 · Rui Feng, Yuang Chen, Yu Hu, Jun Du, Jiahong Yuan arxiv

This letter introduces EGGCodec, a robust neural Encodec framework engineered for electroglottography (EGG) signal reconstruction and F0 extraction. We propose a multi-scale frequency-domain loss function to capture the nuanced relationship between original and reconstructed EGG signals, complemented by a time-domain correlation loss to improve generalization and accuracy. Unlike conventional Encodec models that extract F0 directly from features, EGGCodec leverages reconstructed EGG signals, which more closely correspond to F0. By removing the conventional GAN discriminator, we streamline EGGCodec's training process without compromising efficiency, incurring only negligible performance degradation. Trained on a widely used EGG-inclusive dataset, extensive evaluations demonstrate that EGGCodec outperforms state-of-the-art F0 extraction schemes, reducing mean absolute error (MAE) from 14.14 Hz to 13.69 Hz, and improving voicing decision error (VDE) by 38.2\%. Moreover, extensive ablation experiments validate the contribution of each component of EGGCodec.

📄 PDF Abstract BibTeX arXiv:2508.08924

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EnCodecMAE: Leveraging neural codecs for universal audio representation learning

2023-09-14 · Leonardo Pepino, Pablo Riera, Luciana Ferrer

The goal of universal audio representation learning is to obtain foundational models that can be used for a variety of downstream tasks involving speech, music and environmental sounds. To approach this problem, methods …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation LearningSelf-Supervised Learning+2

SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis

2024-09-11 · Helin Wang, Meng Yu, Jiarui Hai, Chen Chen 외

In this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot textbased speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer deco…

DecoderSpeech Synthesistext-to-speechText to Speech+1

LAV: Audio-Driven Dynamic Visual Generation with Neural Compression and StyleGAN2

2025-05-15 · Jongmin Jung, Dasaem Jeong

This paper introduces LAV (Latent Audio-Visual), a system that integrates EnCodec's neural audio compression with StyleGAN2's generative capabilities to produce visually dynamic outputs driven by pre-recorded audio. Unli…

Audio Compression

Enhancing Suno's Bark Text-to-Speech Model: Addressing Limitations Through Meta's Encodec and Pre-Trained Hubert

2023-04-18 · Social Science Research Network (SSRN) 2023 4 · Devin Schumacher, Francis LaBounty Jr.

Bark, a transformer-based text-to-audio model by Suno, generates highly realistic, multilingual speech as well as other audio, including music, background noise, and simple sound effects. While this model has shown promi…

Audio GenerationExpressive Speech SynthesisSpeech Synthesistext-to-speech+3

Leveraging Pre-Trained Autoencoders for Interpretable Prototype Learning of Music Audio

2024-02-14 · Pablo Alonso-Jiménez, Leonardo Pepino, Roser Batlle-Roca, Pablo Zinemanas 외

We present PECMAE, an interpretable model for music audio classification based on prototype learning. Our model is based on a previous method, APNet, which jointly learns an autoencoder and a prototypical network. Instea…

Audio ClassificationDecoder