paper-with-me

홈 › Papers

Probing the Robustness Properties of Neural Speech Codecs

2025-05-30 · Wei-Cheng Tseng, David Harwath

Neural speech codecs have revolutionized speech coding, achieving higher compression while preserving audio fidelity. Beyond compression, they have emerged as tokenization strategies, enabling language modeling on speech and driving paradigm shifts across various speech processing tasks. Despite these advancements, their robustness in noisy environments remains underexplored, raising concerns about their generalization to real-world scenarios. In this work, we systematically evaluate neural speech codecs under various noise conditions, revealing non-trivial differences in their robustness. We further examine their linearity properties, uncovering non-linear distortions which partly explain observed variations in robustness. Lastly, we analyze their frequency response to identify factors affecting audio fidelity. Our findings provide critical insights into codec behavior and future codec design, as well as emphasizing the importance of noise robustness for their real-world integration.

📄 PDF Abstract BibTeX arXiv:2505.24248

Code (1)

raytzeng/codec-noise-robustness 공식 구현

Tasks

Language ModelingLanguage Modelling

Similar Papers 제목 키워드 기반

Streaming Neural Speech Codecs through Time-Invariant Representations

2026-07-06 · Kélian Estève, Salima Mhdaffar, Mickael Rouvier, Richard Dufour 외 arxiv

Neural speech codecs are increasingly used as intermediate representations in codec-based speech generation systems. TiCodec introduces a factorized representation that separates time-varying speech content from time-inv…

Probing Low Frame Rate Degradation in Neural Audio Codecs

2026-06-15 · Alex Gichamba, Moise Busogi arxiv

Low frame rates in neural audio codecs are attractive for autoregressive speech synthesis, where the generation cost scales linearly with the sequence length. Recent work has demonstrated that codecs can operate at 12.5 …

Speech Synthesis

Analysing the Language of Neural Audio Codecs

2025-09-01 · Joonyong Park, Shinnosuke Takamichi, David M. Chan, Shunsuke Kando 외 arxiv

This study presents a comparative analysis of the statistical and linguistic properties of neural audio codecs (NACs). We investigate discrete speech tokens produced by various NAC models, examining their adherence to li…

Speech Recognition

Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis

2024-09-20 · Lauri Juvela, Xin Wang

Automatic detection of synthetic speech is becoming increasingly important as current synthesis methods are both near indistinguishable from human speech and widely accessible to the public. Audio watermarking and other …

Face SwappingSpeech Synthesis

Conditional probing: measuring usable information beyond a baseline

2021-09-19 · EMNLP 2021 11 · John Hewitt, Kawin Ethayarajh, Percy Liang, Christopher D. Manning

Probing experiments investigate the extent to which neural representations make properties -- like part-of-speech -- predictable. One suggests that a representation encodes a property if probing that representation produ…

Word Embeddings