paper-with-me

Papers

Stage-adaptive audio diffusion modeling

2026-05-06 · Xuanhao Zhang, Chang Li arxiv

Recent progress in diffusion-based audio generation and restoration has substantially improved performance across heterogeneous conditioning regimes, including text-conditioned audio generation and audio-conditioned super-resolution. However, training audio diffusion models remains computationally expensive, and most existing pipelines still rely on static optimization recipes that treat the relative importance of training signals as fixed throughout learning. In this work, we argue that a major source of inefficiency lies in the evolving balance between semantic acquisition and generation-oriented refinement. Early training places stronger emphasis on acquiring condition-aligned semantic structure and coarse global organization, whereas later training increasingly emphasizes temporal consistency, perceptual fidelity, and fine-detail refinement. To characterize this evolving balance, we introduce a progress-based regime variable derived from the training-time slope of an SSL-space discrepancy, which measures semantic progress during training. Based on this signal, we develop three complementary stage-aware mechanisms: decayed SSL guidance for early semantic bootstrapping, self-adaptive timestep sampling driven by the regime variable, and structure-aware regularization activated from convergent grouped organization in parameter space. We evaluate these mechanisms on text-conditioned audio generation and audio-conditioned super-resolution. Across both settings, the proposed stage-aware strategies improve convergence behavior and yield gains on the primary generation and spectral reconstruction metrics over standard static baselines. These results support the view that efficient audio diffusion training can benefit from treating external guidance, internal organization, and optimization emphasis as stage-dependent components rather than fixed ingredients.

📄 PDF Abstract BibTeX arXiv:2605.04547

Code (0)

등록된 구현이 없습니다.

Tasks

Spectral ReconstructionAudio Generation

Similar Papers 제목 키워드 기반

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation

2026-06-10 · Zeyue Tian, Lei Ke, Zhaoyang Liu, Ruibin Yuan 외 arxiv

Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a unified multimodal modeling framework, 2) large-scale, high-quality training d…

Text-to-Music GenerationAudio Generation

RFM-Editing 2: Text-Guided Audio Editing with Rectified Flow Matching and Coarse-to-Fine Diffusion Transformers

2026-06-18 · Liting Gao, Yonggang Zhu, Yaru Chen, Dongyu Wang 외 arxiv

Audio editing aims to modify specific content in an existing audio clip according to a text instruction or description while preserving the remaining acoustic content. Despite the remarkable progress of diffusion models,…

ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling

2025-10-10 · Yuxuan Jiang, Zehua Chen, Zeqian Ju, Yusheng Dai 외 arxiv

Text-to-audio (TTA) generation with fine-grained control signals, e.g., precise timing control or intelligible speech content, has been explored in recent works. However, constrained by data scarcity, their generation pe…

Multi-Task LearningAudio Generation

ViolinDiff: Enhancing Expressive Violin Synthesis with Pitch Bend Conditioning

2024-09-19 · Daewoong Kim, Hao-Wen Dong, Dasaem Jeong

Modeling the natural contour of fundamental frequency (F0) plays a critical role in music audio synthesis. However, transcribing and managing multiple F0 contours in polyphonic music is challenging, and explicit F0 conto…

Audio Synthesis

YingVideo-MV: Music-Driven Multi-Stage Video Generation

2025-12-02 · Jiahui Chen, Weida Wang, Runhua Shi, Huan Yang 외 arxiv

While diffusion model for audio-driven avatar video generation have achieved notable process in synthesizing long sequences with natural audio-visual synchronization and identity consistency, the generation of music-perf…

Video Generation