paper-with-me

홈 › Papers

AVControl: Efficient Framework for Training Audio-Visual Controls

2026-03-25 · Matan Ben-Yosef, Tavi Halperin, Naomi Ken Korem, Mohammad Salama, Harel Cain, Asaf Joseph, Anthony Chen, Urska Jelercic, Ofir Bibi arxiv

Controlling video and audio generation requires diverse modalities, from depth and pose to camera trajectories and audio transformations, yet existing approaches either train a single monolithic model for a fixed set of controls or introduce costly architectural changes for each new modality. We introduce AVControl, a lightweight, extendable framework built on LTX-2, a joint audio-visual foundation model, where each control modality is trained as a separate LoRA on a parallel canvas that provides the reference signal as additional tokens in the attention layers, requiring no architectural changes beyond the LoRA adapters themselves. We show that simply extending image-based in-context methods to video fails for structural control, and that our parallel canvas approach resolves this. On the VACE Benchmark, we outperform all evaluated baselines on depth- and pose-guided generation, inpainting, and outpainting, and show competitive results on camera control and audio-visual benchmarks. Our framework supports a diverse set of independently trained modalities: spatially-aligned controls such as depth, pose, and edges, camera trajectory with intrinsics, sparse motion control, video editing, and, to our knowledge, the first modular audio-visual controls for a joint generation model. Our method is both compute- and data-efficient: each modality requires only a small dataset and converges within a few hundred to a few thousand training steps, a fraction of the budget of monolithic alternatives. We publicly release our code and trained LoRA checkpoints.

📄 PDF Abstract BibTeX arXiv:2603.24793

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Generation

Similar Papers 제목 키워드 기반

Closing the Navigation Compliance Gap in End-to-end Autonomous Driving

2025-12-11 · Hanfeng Wu, Marlon Steiner, Michael Schmidt, Alvaro Marcos-Ramiro 외 arxiv

Trajectory-scoring planners achieve high navigation compliance when following the expert's original command, yet they struggle at intersections when presented with alternative commands; over 30 percent of such commands a…

Autonomous Driving

Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information

2025-10-07 · Christian Marinoni, Riccardo Fosco Gramaccioni, Eleonora Grassucci, Danilo Comminiello arxiv

The generation of sounding videos has seen significant advancements with the advent of diffusion models. However, existing methods often lack the fine-grained control needed to generate viewpoint-specific content from la…

Music ControlNet: A model similar to SD ControlNetD that can accurately control music generation

2023-11-07 · . 2023 11 · Wu, Shih-Lun and Donahue, Chris and Watanabe, Shinji and Bryan 외

Text-to-music generation models are now capable of generating high-quality music audio in broad styles. However, text control is primarily suitable for the manipulation of global musical attributes like genre, mood, and …

Music GenerationRhythmText-to-Music Generation

LiLAC: A Lightweight Latent ControlNet for Musical Audio Generation

2025-06-13 · Tom Baker, Javier Nistal

Text-to-audio diffusion models produce high-quality and diverse music but many, if not most, of the SOTA models lack the fine-grained, time-varying controls essential for music production. ControlNet enables attaching ex…

Audio Generation

SoundPlot: An Open-Source Framework for Birdsong Acoustic Analysis and Neural Synthesis with Interactive 3D Visualization

2026-01-19 · Naqcho Ali Mehdi, Mohammad Adeel, Aizaz Ali Larik arxiv

We present SoundPlot, an open-source framework for analyzing avian vocalizations through acoustic feature extraction, dimensionality reduction, and neural audio synthesis. The system transforms audio signals into a multi…

Dimensionality Reduction