paper-with-me

홈 › Papers

Speech Tokenizer is Key to Consistent Representation

2025-07-09 · Wonjin Jung, Sungil Kang, Dong-Yeon Cho arxiv

Speech tokenization is crucial in digital speech processing, converting continuous speech signals into discrete units for various computational tasks. This paper introduces a novel speech tokenizer with broad applicability across downstream tasks. While recent advances in residual vector quantization (RVQ) have incorporated semantic elements, they often neglect critical acoustic features. We propose an advanced approach that simultaneously encodes both linguistic and acoustic information, preserving prosodic and emotional content. Our method significantly enhances speech representation fidelity across diverse applications. Empirical evaluations demonstrate its effectiveness in speech coding, voice conversion, emotion recognition, and multimodal language modeling, without requiring additional training. This versatility underscores its potential as a key tool for advancing AI-driven speech processing.

📄 PDF Abstract BibTeX arXiv:2507.06802

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionVoice Conversion

Similar Papers 제목 키워드 기반

SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

2023-08-31 · Xin Zhang, Dong Zhang, ShiMin Li, Yaqian Zhou 외

Current speech large language models build upon discrete speech representations, which can be categorized into semantic tokens and acoustic tokens. However, existing speech tokens are not specifically designed for speech…

DecoderLanguage ModelingLanguage ModellingQuantization+2

LoSATok: Low-dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation

2026-05-27 · Zhisheng Zhang, Xiang Li, Yixuan Zhou, Jing Peng 외 arxiv

Audio tokenizers are fundamental to unifying audio understanding and generation. Understanding requires high-level semantics, while generation demands semantic and acoustic details. Existing unified tokenizers jointly en…

Audio Generation

SpeechLM: Enhanced Speech Pre-Training with Unpaired Textual Data

2022-09-30 · Ziqiang Zhang, Sanyuan Chen, Long Zhou, Yu Wu 외

How to boost speech pre-training with textual data is an unsolved problem due to the fact that speech and text are very different modalities with distinct characteristics. In this paper, we propose a cross-modal Speech a…

Language ModelingLanguage Modellingspeech-recognitionSpeech Recognition

Scaling Speech Tokenizers with Diffusion Autoencoders

2026-02-06 · Yuancheng Wang, Zhenyu Tang, Yun Wang, Arthur Hinsvark 외 arxiv

Speech tokenizers are foundational to speech language models, yet existing approaches face two major challenges: (1) balancing trade-offs between encoding semantics for understanding and acoustics for reconstruction, and…

Continuous Speech Tokenizer in Text To Speech

2024-10-22 · Yixing Li, Ruobing Xie, Xingwu Sun, Yu Cheng 외

The fusion of speech and language in the era of large language models has garnered significant attention. Discrete speech token is often utilized in text-to-speech tasks for speech compression and portability, which is c…

Language ModelingLanguage Modellingtext-to-speechText to Speech