paper-with-me

홈 › Papers

Stack-and-Delay: a new codebook pattern for music generation

2023-09-15 · Gael Le Lan, Varun Nagaraja, Ernie Chang, David Kant, Zhaoheng Ni, Yangyang Shi, Forrest Iandola, Vikas Chandra

In language modeling based music generation, a generated waveform is represented by a sequence of hierarchical token stacks that can be decoded either in an auto-regressive manner or in parallel, depending on the codebook patterns. In particular, flattening the codebooks represents the highest quality decoding strategy, while being notoriously slow. To this end, we propose a novel stack-and-delay style of decoding strategy to improve upon the flat pattern decoding where generation speed is four times faster as opposed to vanilla flat decoding. This brings the inference time close to that of the delay decoding strategy, and allows for faster inference on GPU for small batch sizes. For the same inference efficiency budget as the delay pattern, we show that the proposed approach performs better in objective evaluations, almost closing the gap with the flat pattern in terms of quality. The results are corroborated by subjective evaluations which show that samples generated by the new model are slightly more often preferred to samples generated by the competing model given the same text prompts.

📄 PDF Abstract BibTeX arXiv:2309.08804

Code (0)

등록된 구현이 없습니다.

Tasks

GPULanguage ModelingLanguage ModellingMusic Generation

Methods 이 논문이 사용한 방법론

SPEED The monocular depth estimation (MDE) is the task of estimating depth from a single frame. This information is an essential knowledge in many computer vision tasks such as scene…

Similar Papers 제목 키워드 기반

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations

2025-07-16 · Yichen Han, Xiaoyang Hao, Keming Chen, Weibo Xiong 외 arxiv

Text-to-speech (TTS) synthesis has seen renewed progress under the discrete modeling paradigm. Existing autoregressive approaches often rely on single-codebook representations, which suffer from significant information l…

Audio Generation

An Independence-promoting Loss for Music Generation with Language Models

2024-06-04 · Jean-Marie Lemercier, Simon Rouard, Jade Copet, Yossi Adi 외

Music generation schemes using language modeling rely on a vocabulary of audio tokens, generally provided as codes in a discrete latent space learnt by an auto-encoder. Multi-stage quantizers are often employed to produc…

Language ModelingLanguage ModellingMusic Generation

Modeling Music as a Time-Frequency Image: A 2D Tokenizer for Music Generation

2026-05-15 · Yuqing Cheng, Xingyu Ma, Guochen Yu, Xiaotao Gu arxiv

Autoregressive music generation depends strongly on the audio tokenizer. Existing high-fidelity codecs often use residual multi-codebook quantization, which preserves reconstruction quality but complicates language model…

Music Generation

UniMuMo: Unified Text, Music and Motion Generation

2024-10-06 · Han Yang, Kun Su, Yutong Zhang, Jiaben Chen 외

We introduce UniMuMo, a unified multimodal model capable of taking arbitrary text, music, and motion data as input conditions to generate outputs across all three modalities. To address the lack of time-synchronized data…

DecoderMotion Generation

TokenDance: Token-to-Token Music-to-Dance Generation with Bidirectional Mamba

2026-03-28 · Ziyue Yang, Kaixing Yang, Xulong Tang arxiv

Music-to-dance generation has broad applications in virtual reality, dance education, and digital character animation. However, the limited coverage of existing 3D dance datasets confines current models to a narrow subse…

Motion Synthesis