paper-with-me

Papers

Unlocking Temporal Flexibility: Neural Speech Codec with Variable Frame Rate

2025-05-22 · Hanglei Zhang, Yiwei Guo, Zhihan Li, Xiang Hao, Xie Chen, Kai Yu

Most neural speech codecs achieve bitrate adjustment through intra-frame mechanisms, such as codebook dropout, at a Constant Frame Rate (CFR). However, speech segments inherently have time-varying information density (e.g., silent intervals versus voiced regions). This property makes CFR not optimal in terms of bitrate and token sequence length, hindering efficiency in real-time applications. In this work, we propose a Temporally Flexible Coding (TFC) technique, introducing variable frame rate (VFR) into neural speech codecs for the first time. TFC enables seamlessly tunable average frame rates and dynamically allocates frame rates based on temporal entropy. Experimental results show that a codec with TFC achieves optimal reconstruction quality with high flexibility, and maintains competitive performance even at lower frame rates. Our approach is promising for the integration with other efforts to develop low-frame-rate neural speech codecs for more efficient downstream tasks.

📄 PDF Abstract BibTeX arXiv:2505.16845

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

CodecSlime: Temporal Redundancy Compression of Neural Speech Codec via Dynamic Frame Rate

2025-06-26 · Hankun Wang, Yiwei Guo, Chongtian Shao, Bohan Li 외

Neural speech codecs have been widely used in audio compression and various downstream tasks. Current mainstream codecs are fixed-frame-rate (FFR), which allocate the same number of tokens to every equal-duration slice. …

Audio Compression

UBGAN: Enhancing Coded Speech with Blind and Guided Bandwidth Extension

2025-05-22 · Kishan Gupta, Srikanth Korse, Andreas Brendel, Nicola Pia 외

In practical application of speech codecs, a multitude of factors such as the quality of the radio connection, limiting hardware or required user experience necessitate trade-offs between achievable perceptual quality, e…

Bandwidth ExtensionGenerative Adversarial Network

Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation

2024-06-11 · Hanzhao Li, Liumeng Xue, Haohan Guo, Xinfa Zhu 외

The multi-codebook speech codec enables the application of large language models (LLM) in TTS but bottlenecks efficiency and robustness due to multi-sequence prediction. To avoid this obstacle, we propose Single-Codec, a…

FuseCodec: Semantic-Contextual Fusion and Supervision for Neural Codecs

2025-09-14 · Md Mubtasim Ahasan, Rafat Hasan Khan, Tasnim Mohiuddin, Aman Chadha 외 arxiv

Speech tokenization enables discrete representation and facilitates speech language modeling. However, existing neural codecs capture low-level acoustic features, overlooking the semantic and contextual cues inherent to …

Representation LearningSpeech Synthesis

Exploring the Rate-Distortion-Complexity Optimization in Neural Image Compression

2023-05-12 · Yixin Gao, Runsen Feng, Zongyu Guo, Zhibo Chen

Despite a short history, neural image codecs have been shown to surpass classical image codecs in terms of rate-distortion performance. However, most of them suffer from significantly longer decoding times, which hinders…

Image Compression