paper-with-me

Papers Audio Generation

“Audio Generation” 태그가 달린 논문 392편 · 필터 해제

StepAudio 3 Gen Technical Report

2026-09-11 · Bin Lin, Bo Zhao, Boyang Wang, Boyang Zhang 외 hf

We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound effects, music, vibe speech, and mixtures of multiple audio types…

Audio Generation

PRISM-Bench: An Audio-Centric Diagnostic Benchmark for Text-to-Audio-Video Generation

2026-09-04 · Yuchen Sun, Qian Yang, Jun Wang, Detai Xin 외 arxiv

Text-to-audio-video (T2AV) generation has advanced rapidly, but its evaluation still underestimates the audio modality. Existing benchmarks either treat audio as an auxiliary component of video quality or assess it in is…

Audio GenerationVideo Generation

The Attention Triangle in Audio-Video Models

2026-09-03 · Sagi Polaczek, Noa Kraicer, Gal Metzer, Zhuo Ning 외 hf

Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and…

Audio GenerationVideo Generation

The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation

2026-09-02 · Yichen Liu, Quanwei Zhang, Haozhe Wang, Donghao Zhou 외 hf

Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization. However, they still provide limited control over when shot transitions occur and dialogue is spoken. …

Audio GenerationVideo Generation

TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation

2026-08-25 · Xiaoda Yang, Yuxiang Liu, Kaiwen Zheng, Yuan Liu 외 arxiv

Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal computation of large models make inference prohibitively expensive. …

Audio Generation

Geometric Iterative Retrieval for Neural Audio Codec Resynthesis

2026-08-19 · Leo Schmidt-Traub, Frédéric Berdoz, Luca A. Lanzendörfer, Roger Wattenhofer arxiv

Neural audio codecs based on Residual Vector Quantization (RVQ) have become the dominant discrete representation for token-based general audio generation, yet resynthesizing high-quality audio from coarse codec tokens re…

Audio Generation

VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation

2026-08-19 · Yinming Huang, Shuyuan Tu, Xi Yan, Zihan Yang 외 arxiv

Using reinforcement learning to post-train joint video-audio generation models requires a reward signal. Existing methods construct this reward by combining metrics for individual quality dimensions, including audio qual…

Reinforcement LearningVideo GenerationAudio Generation

MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation

2026-08-10 · Shuyu Li, Kejun Zhang, Jiahe Lei, Shulei Ji 외 arxiv

Text-to-music generation has advanced rapidly, but current systems still rely primarily on global text prompts, leaving the structural organization of generated music implicit and difficult to inspect, control, or revise…

Text-to-Music GenerationAudio Generation

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

2026-08-03 · Yu Zhang, Ruiqi Li, Changhao Pan, Ke Lei 외 hf

Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference re…

Audio Generation

Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds

2026-08-01 · Masaki Yoshida, Ren Togo, Takahiro Ogawa, Miki Haseyama arxiv

3D Gaussian Splatting (3DGS) turns captured or generated imagery into photorealistic 3D world simulations that users can freely explore, yet these worlds remain silent. Because existing audio generation methods condition…

Audio Generation

FlashDiff: Efficient Regional Execution and Scheduling for Diffusion Model Serving

2026-07-13 · Yaqi Qiao, Ping He, Songrun Xie, Ayush Barik 외 arxiv

Diffusion models have become the central backbone for modern image, video, and audio generation, but their efficient service remains a challenge. Unlike autoregressive decoding, diffusion inference repeatedly updates hig…

Audio Generation

A Quantized Native Runtime for On-Device Semantic Audio Generation

2026-07-09 · Matteo Spanio, Antonio Rodà hf

Semantic audio applications increasingly require controllable generation on commodity and embedded hardware rather than through framework-heavy datacenter stacks. We present aria, a dependency-free native runtime that ru…

Audio Generation

Unified Audio Intelligence Without Regressing on Text Intelligence

2026-07-06 · Zhifeng Kong, Sang-gil Lee, Jaehyeon Kim, Boxin Wang 외 arxiv

Audio intelligence involves understanding, reasoning about, and generating both audio and speech. In this work, we introduce Nemotron-Labs-Audex-30B-A3B (Audex), a unified audio-text LLM built on Nemotron-Cascade-2-30B-A…

multimodal generationSpeech RecognitionAudio Generation

Synthetic Audio Generation Framework for Air Traffic Control Speech Recognition

2026-06-19 · Raphaël Bagat, Zhe Zhang, Junichi Yamagishi, Irina Illina 외 arxiv

Automatic Speech Recognition (ASR) systems, despite achieving remarkable accuracy in general-purpose domains with native speech (L1), struggle in domains like Air Traffic Control (ATC) due to strong channel noise, a pres…

Synthetic Data GenerationSpeech RecognitionAudio GenerationVoice Conversion

FoleyGenEx: Unified Video-to-Audio Generation with Multi-Modal Control, Temporal Alignment, and Semantic Precision

2026-06-12 · Shiyao Wang, Xijuan Zeng, Hui Wang, Shiwan Zhao 외 arxiv

We present FoleyGenEx, a unified video-to-audio (VTA) framework integrating multi-modal control, frame-level temporal alignment, and fine-grained semantics, enabling synchronized, versatile audio synthesis for diverse ta…

Data AugmentationAudio Generation

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation

2026-06-10 · Zeyue Tian, Lei Ke, Zhaoyang Liu, Ruibin Yuan 외 arxiv

Audio and music generation based on flexible multimodal control signals is a widely applicable topic, with the following key challenges: 1) a unified multimodal modeling framework, 2) large-scale, high-quality training d…

Text-to-Music GenerationAudio Generation

AnyAudio-Judge: A Dynamic Rubric-Based Benchmark and Evaluator for Audio Instruction Following

2026-06-02 · Haitao Li, Tian Tan, Yuguang Yang, Shan Yang 외 arxiv

The rapid advancement of instruction-guided audio generation has highlighted the critical need for robust alignment evaluation. Current automated evaluation methods heavily rely on holistic scoring from general-purpose l…

Reinforcement LearningInstruction FollowingAudio Generation

SketchSong: Hierarchical Song Generation with Sketch Planning and Fine-Grained Multi-Track Modeling

2026-06-02 · Xiaoyue Duan, Nanxing Hu, Yutang Feng, Xudong Yan 외 arxiv

Recent song generation systems can synthesize realistic audio, yet generating complete songs remains challenging for two reasons. First, explicit song-level arrangement planning remains limited in existing methods, so mo…

Audio Generation

JenBridge: Adaptive Long-Form Video Soundtracking across Scene Transitions

2026-06-01 · Jiashuo Yu, Yao Yao, Boyu Chen, Alex Wang arxiv

We address the challenge of generating high-fidelity, long-form soundtracks that remain coherent across scene transitions. Existing AI music systems are mainly designed for short, isolated clips and lack mechanisms to en…

Audio Generation

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment

2026-05-29 · Jun-Hak Yun, Seung-Bin Kim, Seong-Whan Lee arxiv

Recent advancements in text-guided audio generation have yielded promising results in diverse domains, including sound effects, speech, and music. However, jointly generating speech with environmental audio remains chall…

Audio Generation
1–20 / 392 다음 →