paper-with-me

홈 › Papers

Language Model Based Text-to-Audio Generation: Anti-Causally Aligned Collaborative Residual Transformers

2025-10-06 · Juncheng Wang, Chao Xu, Cheng Yu, Zhe Hu, Haoyu Xie, Guoqi Yu, Lei Shang, Shujun Wang arxiv

While language models (LMs) paired with residual vector quantization (RVQ) tokenizers have shown promise in text-to-audio (T2A) generation, they still lag behind diffusion-based models by a non-trivial margin. We identify a critical dilemma underpinning this gap: incorporating more RVQ layers improves audio reconstruction fidelity but exceeds the generation capacity of conventional LMs. To address this, we first analyze RVQ dynamics and uncover two key limitations: 1) orthogonality of features across RVQ layers hinders effective LMs training, and 2) descending semantic richness in tokens from deeper RVQ layers exacerbates exposure bias during autoregressive decoding. Based on these insights, we propose Siren, a novel LM-based framework that employs multiple isolated transformers with causal conditioning and anti-causal alignment via reinforcement learning. Extensive experiments demonstrate that Siren outperforms both existing LM-based and diffusion-based T2A systems, achieving state-of-the-art results. By bridging the representational strengths of LMs with the fidelity demands of audio synthesis, our approach repositions LMs as competitive contenders against diffusion models in T2A tasks. Moreover, by aligning audio representations with linguistic structures, Siren facilitates a promising pathway toward unified multi-modal generation frameworks.

📄 PDF Abstract BibTeX arXiv:2510.04577

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement LearningAudio Generation

Similar Papers 제목 키워드 기반

LiveBand: Live Accompaniment Generation in the Audio Domain

2026-06-02 · Marco Pasini, Javier Nistal, Ben Hayes, Mathias Rose Bjare 외 arxiv

We present LiveBand, a real-time system that generates high-fidelity music accompaniments to live audio input, respecting strict causal constraints. Our method trains a causal transformer generator in the continuous late…

FoleyBench: A Benchmark For Video-to-Audio Models

2025-11-17 · Satvik Dixit, Koichi Saito, Zhi Zhong, Yuki Mitsufuji 외 arxiv

Video-to-audio generation (V2A) is of increasing importance in domains such as film post-production, AR/VR, and sound design, particularly for the creation of Foley sound effects synchronized with on-screen actions. Fole…

Audio GenerationVideo Alignment

Neuron-Level Emotion Control in Speech-Generative Large Audio-Language Models

2026-03-18 · Xiutian Zhao, Ismail Rasim Ulgen, Philipp Koehn, Björn Schuller 외 arxiv

Large audio-language models (LALMs) can produce expressive speech, yet reliable emotion control remains elusive: conversions often miss the target affect and may degrade linguistic fidelity through refusals, hallucinatio…

CaM-Gen: Causally-aware Guided Text Generation

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Content is created for a well-defined purpose, often described by a metric or signal represented in the form of structured information. The relationship between the goal (metrics) of target content and the content itself…

Causal InferenceText Generation

CaM-Gen:Causally-aware Metric-guided Text Generation

2020-10-24 · Navita Goyal, Roodram Paneri, Ayush Agarwal, Udit Kalani 외

Content is created for a well-defined purpose, often described by a metric or signal represented in the form of structured information. The relationship between the goal (metrics) of target content and the content itself…

Causal InferenceText Generation