paper-with-me

Papers Speech Tokenization

“Speech Tokenization” 태그가 달린 논문 21편 · 필터 해제

LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization

2025-06-20 · DaeJin Jo, Jeeyoung Yun, Byungseok Roh, Sungwoong Kim

With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text, enabling unified modeling across modalities. Recent speech tokenization approache…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+6

Factorized RVQ-GAN For Disentangled Speech Tokenization

2025-06-18 · Sameer Khurana, Dominik Klement, Antoine Laurent, Dominik Bobos 외

We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single model. HAC leverages two knowledge dist…

DisentanglementKnowledge DistillationSpeech Tokenization

Audio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework

2025-05-24 · Binhao Ma, Hanqing Guo, Zhengping Jay Luo, Rui Duan

Recent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced the naturalness and flexibility of human computer interaction by enabling seamless understanding across text, vision, and audio moda…

Adversarial AttackSpeech TokenizationSpeech-to-TextSpeech-to-Text Translation

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models

2025-05-23 · Shunsuke Kando, Yusuke Miyao, Shinnosuke Takamichi

The purpose of speech tokenization is to transform a speech signal into a sequence of discrete representations, serving as the foundation for speech language models (SLMs). While speech tokenization has many options, the…

Speech TokenizationSpoken Language Understanding

Impact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English

2025-05-20 · Haoyang Zhang, Hexin Liu, Xiangyu Zhang, Qiquan Zhang 외

The speech tokenizer plays a crucial role in recent speech tasks, generally serving as a bridge between speech signals and language models. While low-frame-rate codecs are widely employed as speech tokenizers, the impact…

Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSpeech Tokenization+2

TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling

2025-04-09 · Liang-Hsuan Tseng, Yi-Chang Chen, Kuan-Yi Lee, Da-Shan Shiu 외

Recent efforts target spoken language models (SLMs) that not only listen but also speak for more natural human-LLM interaction. Joint speech-text modeling is a promising direction to achieve this. However, the effectiven…

Language ModelingLanguage Modellingparameter-efficient fine-tuningSpeech Tokenization

UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation

2025-03-02 · Alexander H. Liu, Sang-gil Lee, Chao-Han Huck Yang, Yuan Gong 외

Pre-training and representation learning have been playing an increasingly important role in modern speech processing. Nevertheless, different applications have been relying on different foundation models, since predomin…

DecoderRepresentation Learningspeech-recognitionSpeech Recognition+4

Recent Advances in Discrete Speech Tokens: A Review

2025-02-10 · Yiwei Guo, Zhihan Li, Hankun Wang, Bohan Li 외

The rapid advancement of speech generation technologies in the era of large language models (LLMs) has established discrete speech tokens as a foundational paradigm for speech representation. These tokens, characterized …

Language ModelingLanguage ModellingSpeech Tokenization

BEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection

2024-11-21 · Anup Singh, Kris Demuynck, Vipul Arora

Spoken term detection (STD) is often hindered by reliance on frame-level features and the computationally intensive DTW-based template matching, limiting its practicality. To address these challenges, we propose a novel …

MambaSelf-Supervised LearningSpeech TokenizationTemplate Matching

DC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models

2024-10-31 · Heng-Jui Chang, Hongyu Gong, Changhan Wang, James Glass 외

Spoken language models (SLMs) have gained increasing attention with advancements in text-based, decoder-only language models. SLMs process text and speech, enabling simultaneous speech understanding and generation. This …

DecoderResynthesisSpeech Tokenization

DM-Codec: Distilling Multimodal Representations for Speech Tokenization

2024-10-19 · Md Mubtasim Ahasan, Md Fahim, Tasnim Mohiuddin, A K M Mahbubur Rahman 외

Recent advancements in speech-language models have yielded significant improvements in speech tokenization and synthesis. However, effectively mapping the complex, multidimensional attributes of speech into discrete toke…

Self-Supervised LearningSpeech Tokenization

Sylber: Syllabic Embedding Representation of Speech from Raw Audio

2024-10-09 · Cheol Jun Cho, Nicholas Lee, Akshat Gupta, Dhruv Agarwal 외

Syllables are compositional units of spoken language that efficiently structure human speech perception and production. However, current neural speech representations lack such structure, resulting in dense token sequenc…

Language ModelingLanguage ModellingSelf-Supervised LearningSpeech Tokenization

SyllableLM: Learning Coarse Semantic Units for Speech Language Models

2024-10-05 · Alan Baade, Puyuan Peng, David Harwath

Language models require tokenized inputs. However, tokenization strategies for continuous data like audio and vision are often based on simple heuristics such as fixed sized convolutions or discrete clustering, which do …

ClusteringLanguage ModelingLanguage ModellingSpeech Tokenization+1

Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT

2024-09-16 · Ryota Komatsu, Takahiro Shinozaki

Self-supervised speech representation learning has become essential for extracting meaningful features from untranscribed audio. Recent advances highlight the potential of deriving discrete symbols from the features corr…

Acoustic Unit DiscoveryClusteringData AugmentationRepresentation Learning+3

LAST: Language Model Aware Speech Tokenization

2024-09-05 · Arnon Turetzky, Yossi Adi

Speech tokenization serves as the foundation of speech language model (LM), enabling them to perform various tasks such as spoken language modeling, text-to-speech, speech-to-text, etc. Most speech tokenizers are trained…

Language ModelingLanguage ModellingmodelQuantization+4

STAB: Speech Tokenizer Assessment Benchmark

2024-09-04 · Shikhar Vashishth, Harman Singh, Shikhar Bharadwaj, Sriram Ganapathy 외

Representing speech as discrete tokens provides a framework for transforming speech into a format that closely resembles text, thus enabling the use of speech as an input to the widely successful large language models (L…

Speech Tokenization

dMel: Speech Tokenization made Simple

2024-07-22 · Richard He Bai, Tatiana Likhomanenko, Ruixiang Zhang, Zijin Gu 외

Large language models have revolutionized natural language processing by leveraging self-supervised pretraining on vast textual data. Inspired by this success, researchers have investigated various compression-based spee…

DecoderLanguage ModelingLanguage Modellingspeech-recognition+3

Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing

2024-06-04 · Viet Anh Trinh, Rosy Southwell, Yiwen Guan, Xinlu He 외

Recent work on discrete speech tokenization has paved the way for models that can seamlessly perform multiple tasks across modalities, e.g., speech recognition, text to speech, speech to speech translation. Moreover, lar…

DecoderLanguage ModelingLanguage ModellingLarge Language Model+6

Scaling Properties of Speech Language Models

2024-03-31 · Santiago Cuervo, Ricard Marxer

Speech Language Models (SLMs) aim to learn language from raw audio, without textual resources. Despite significant advances, our current models exhibit weak syntax and semantic abilities. However, if the scaling properti…

Speech Tokenization

BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

2024-02-12 · Mateusz Łajszczak, Guillermo Cámbara, Yang Li, Fatih Beyhan 외

We introduce a text-to-speech (TTS) model called BASE TTS, which stands for $\textbf{B}$ig $\textbf{A}$daptive $\textbf{S}$treamable TTS with $\textbf{E}$mergent abilities. BASE TTS is the largest TTS model to-date, trai…

DecoderDisentanglementSpeech Tokenizationtext-to-speech+1
1–20 / 21 다음 →