Papers Speech Tokenization
“Speech Tokenization” 태그가 달린 논문 21편 · 필터 해제
LM-SPT: LM-Aligned Semantic Distillation for Speech Tokenization
With the rapid progress of speech language models (SLMs), discrete speech tokens have emerged as a core interface between speech and text, enabling unified modeling across modalities. Recent speech tokenization approache…
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Language ModelingLanguage Modelling+6Factorized RVQ-GAN For Disentangled Speech Tokenization
We propose Hierarchical Audio Codec (HAC), a unified neural speech codec that factorizes its bottleneck into three linguistic levels-acoustic, phonetic, and lexical-within a single model. HAC leverages two knowledge dist…
DisentanglementKnowledge DistillationSpeech TokenizationAudio Jailbreak Attacks: Exposing Vulnerabilities in SpeechGPT in a White-Box Framework
Recent advances in Multimodal Large Language Models (MLLMs) have significantly enhanced the naturalness and flexibility of human computer interaction by enabling seamless understanding across text, vision, and audio moda…
Adversarial AttackSpeech TokenizationSpeech-to-TextSpeech-to-Text TranslationExploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
The purpose of speech tokenization is to transform a speech signal into a sequence of discrete representations, serving as the foundation for speech language models (SLMs). While speech tokenization has many options, the…
Speech TokenizationSpoken Language UnderstandingImpact of Frame Rates on Speech Tokenizer: A Case Study on Mandarin and English
The speech tokenizer plays a crucial role in recent speech tasks, generally serving as a bridge between speech signals and language models. While low-frame-rate codecs are widely employed as speech tokenizers, the impact…
Automatic Speech Recognitionspeech-recognitionSpeech RecognitionSpeech Tokenization+2TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
Recent efforts target spoken language models (SLMs) that not only listen but also speak for more natural human-LLM interaction. Joint speech-text modeling is a promising direction to achieve this. However, the effectiven…
Language ModelingLanguage Modellingparameter-efficient fine-tuningSpeech TokenizationUniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
Pre-training and representation learning have been playing an increasingly important role in modern speech processing. Nevertheless, different applications have been relying on different foundation models, since predomin…
DecoderRepresentation Learningspeech-recognitionSpeech Recognition+4Recent Advances in Discrete Speech Tokens: A Review
The rapid advancement of speech generation technologies in the era of large language models (LLMs) has established discrete speech tokens as a foundational paradigm for speech representation. These tokens, characterized …
Language ModelingLanguage ModellingSpeech TokenizationBEST-STD: Bidirectional Mamba-Enhanced Speech Tokenization for Spoken Term Detection
Spoken term detection (STD) is often hindered by reliance on frame-level features and the computationally intensive DTW-based template matching, limiting its practicality. To address these challenges, we propose a novel …
MambaSelf-Supervised LearningSpeech TokenizationTemplate MatchingDC-Spin: A Speaker-invariant Speech Tokenizer for Spoken Language Models
Spoken language models (SLMs) have gained increasing attention with advancements in text-based, decoder-only language models. SLMs process text and speech, enabling simultaneous speech understanding and generation. This …
DecoderResynthesisSpeech TokenizationDM-Codec: Distilling Multimodal Representations for Speech Tokenization
Recent advancements in speech-language models have yielded significant improvements in speech tokenization and synthesis. However, effectively mapping the complex, multidimensional attributes of speech into discrete toke…
Self-Supervised LearningSpeech TokenizationSylber: Syllabic Embedding Representation of Speech from Raw Audio
Syllables are compositional units of spoken language that efficiently structure human speech perception and production. However, current neural speech representations lack such structure, resulting in dense token sequenc…
Language ModelingLanguage ModellingSelf-Supervised LearningSpeech TokenizationSyllableLM: Learning Coarse Semantic Units for Speech Language Models
Language models require tokenized inputs. However, tokenization strategies for continuous data like audio and vision are often based on simple heuristics such as fixed sized convolutions or discrete clustering, which do …
ClusteringLanguage ModelingLanguage ModellingSpeech Tokenization+1Self-Supervised Syllable Discovery Based on Speaker-Disentangled HuBERT
Self-supervised speech representation learning has become essential for extracting meaningful features from untranscribed audio. Recent advances highlight the potential of deriving discrete symbols from the features corr…
Acoustic Unit DiscoveryClusteringData AugmentationRepresentation Learning+3LAST: Language Model Aware Speech Tokenization
Speech tokenization serves as the foundation of speech language model (LM), enabling them to perform various tasks such as spoken language modeling, text-to-speech, speech-to-text, etc. Most speech tokenizers are trained…
Language ModelingLanguage ModellingmodelQuantization+4STAB: Speech Tokenizer Assessment Benchmark
Representing speech as discrete tokens provides a framework for transforming speech into a format that closely resembles text, thus enabling the use of speech as an input to the widely successful large language models (L…
Speech TokenizationdMel: Speech Tokenization made Simple
Large language models have revolutionized natural language processing by leveraging self-supervised pretraining on vast textual data. Inspired by this success, researchers have investigated various compression-based spee…
DecoderLanguage ModelingLanguage Modellingspeech-recognition+3Discrete Multimodal Transformers with a Pretrained Large Language Model for Mixed-Supervision Speech Processing
Recent work on discrete speech tokenization has paved the way for models that can seamlessly perform multiple tasks across modalities, e.g., speech recognition, text to speech, speech to speech translation. Moreover, lar…
DecoderLanguage ModelingLanguage ModellingLarge Language Model+6Scaling Properties of Speech Language Models
Speech Language Models (SLMs) aim to learn language from raw audio, without textual resources. Despite significant advances, our current models exhibit weak syntax and semantic abilities. However, if the scaling properti…
Speech TokenizationBASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
We introduce a text-to-speech (TTS) model called BASE TTS, which stands for $\textbf{B}$ig $\textbf{A}$daptive $\textbf{S}$treamable TTS with $\textbf{E}$mergent abilities. BASE TTS is the largest TTS model to-date, trai…
DecoderDisentanglementSpeech Tokenizationtext-to-speech+1