paper-with-me

Papers

dMel: Speech Tokenization made Simple

2024-07-22 · Richard He Bai, Tatiana Likhomanenko, Ruixiang Zhang, Zijin Gu, Zakaria Aldeneh, Navdeep Jaitly

Large language models have revolutionized natural language processing by leveraging self-supervised pretraining on vast textual data. Inspired by this success, researchers have investigated various compression-based speech tokenization methods to discretize continuous speech signals, enabling the application of language modeling techniques to discrete tokens. However, audio compressor introduces additional complexity and computational cost, and often fail on out-of-domain audio signals. In this work, we introduce a novel speech representation (dmel) that discretizes mel-filterbank channels into intensity bins, creating a simpler yet more effective representation compared to existing speech tokenization methods. Our approach demonstrates superior performance in preserving audio content, robustness to out-of-domain data, and offers a training-free, natural, and streamable representation. To address the high-dimensional nature of log-mel spectrograms, we propose an efficient parallel encoding and decoding method for high-dimensional tokens using an LM-style transformer architecture. This innovation enables us to develop RichTTS and RichASR, two models sharing the same architecture while achieving comparable or better results than specialized existing methods. Our results demonstrate the effectiveness of dmel in achieving high performance on both speech synthesis and recognition tasks within a unified framework, paving the way for efficient and effective joint modeling of speech and text.

📄 PDF Abstract BibTeX arXiv:2407.15835

Code (1)

apple/dmel 공식 구현 pytorch

Tasks

DecoderLanguage ModelingLanguage Modellingspeech-recognitionSpeech RecognitionSpeech SynthesisSpeech Tokenization

Similar Papers 제목 키워드 기반

Benchmarking Diverse-Modal Entity Linking with Generative Models

2023-05-27 · Sijia Wang, Alexander Hanbo Li, Henry Zhu, Sheng Zhang 외

Entities can be expressed in diverse formats, such as texts, images, or column names and cell values in tables. While existing entity linking (EL) models work well on per modality configuration, such as text-only EL, vis…

BenchmarkingDecoderEntity LinkingVisual Grounding

AtteSTNet -- An attention and subword tokenization based approach for code-switched text hate speech detection

2021-12-10 · Geet Shingi, Vedangi Wagh, Kishor Wagh, Sharmila Wagh

Recent advancements in technology have led to a boost in social media usage which has ultimately led to large amounts of user-generated data which also includes hateful and offensive speech. The language used in social m…

Hate Speech Detection

ProsodyLM: Uncovering the Emerging Prosody Processing Capabilities in Speech Language Models

2025-07-27 · Kaizhi Qian, Xulin Fan, Junrui Ni, Slava Shechtman 외 arxiv

Speech language models refer to language models with speech processing and understanding capabilities. One key desirable capability for speech language models is the ability to capture the intricate interdependency betwe…

FedMeld: A Model-dispersal Federated Learning Framework for Space-ground Integrated Networks

2024-12-23 · Qian Chen, Xianhao Chen, Kaibin Huang

To bridge the digital divide, the space-ground integrated networks (SGINs), which will be a key component of the six-generation (6G) mobile networks, are expected to deliver artificial intelligence (AI) services to every…

Federated Learning

MindMelody: A Closed-Loop EEG-Driven System for Personalized Music Intervention

2026-05-02 · Yimeng Zhang, Yueru Sun, Haoyu Gu, Zhanpeng Jin arxiv

Driven by the escalating global burden of mental health conditions, music-based interventions have attracted significant attention as a non-invasive, cost-effective modality for emotion regulation and psychological stres…

Music Generation