paper-with-me

Papers

SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos

2025-10-03 · Amir Dellali, Luca A. Lanzendörfer, Florian Grötschla, Roger Wattenhofer arxiv

We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective, enabling audio-conditioned generation and the seamless synthesis of unconstrained length audio sequences. Additionally, by integrating a shortcut loss into our training process, we achieve rapid generation of high-quality audio samples in as few as eight sampling steps, paving the way for near-real-time applications without requiring dedicated fine-tuning or retraining. We demonstrate that SALSA-V significantly outperforms existing state-of-the-art methods in both audiovisual alignment and synchronization with video content in quantitative evaluation and a human listening study. Furthermore, our use of random masking during training enables our model to match spectral characteristics of reference audio samples, broadening its applicability to professional audio synthesis tasks such as Foley generation and sound design.

📄 PDF Abstract BibTeX arXiv:2510.02916

Code (0)

등록된 구현이 없습니다.

Tasks

Audio Generation

Similar Papers 제목 키워드 기반

SALSA: Spatial Cue-Augmented Log-Spectrogram Features for Polyphonic Sound Event Localization and Detection

2021-10-01 · Thi Ngoc Tho Nguyen, Karn N. Watcharasupat, Ngoc Khanh Nguyen, Douglas L. Jones 외

Sound event localization and detection (SELD) consists of two subtasks, which are sound event detection and direction-of-arrival estimation. While sound event detection mainly relies on time-frequency patterns to disting…

Direction of Arrival EstimationEvent DetectionSound Event DetectionSound Event Localization and Detection

SALSA-Lite: A Fast and Effective Feature for Polyphonic Sound Event Localization and Detection with Microphone Arrays

2021-11-16 · Thi Ngoc Tho Nguyen, Douglas L. Jones, Karn N. Watcharasupat, Huy Phan 외

Polyphonic sound event localization and detection (SELD) has many practical applications in acoustic sensing and monitoring. However, the development of real-time SELD has been limited by the demanding computational requ…

Sound Event Localization and Detection

LSALSA: Accelerated Source Separation via Learned Sparse Coding

2018-02-13 · Benjamin Cowen, Apoorva Nandini Saridena, Anna Choromanska

We propose an efficient algorithm for the generalized sparse coding (SC) inference problem. The proposed framework applies to both the single dictionary setting, where each data point is represented as a sparse combinati…

SegSALSA-STR: A convex formulation to supervised hyperspectral image segmentation using hidden fields and structure tensor regularization

2015-04-27 · Filipe Condessa, Jose Bioucas-Dias, Jelena Kovacevic

We present a supervised hyperspectral image segmentation algorithm based on a convex formulation of a marginal maximum a posteriori segmentation with hidden fields and structure tensor regularization: Segmentation via th…

Hyperspectral Image SegmentationImage SegmentationSegmentationSemantic Segmentation

Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models

2026-06-09 · Qingzi Wang, Xiyang Wu, Guangyao Shi, Dianwei Chen 외 arxiv

Safe social navigation requires robots to distinguish people from ordinary obstacles and to react before danger becomes imminent. We show that pretrained Vision-Language-Action (VLA) models already encode pedestrian-obje…

Collision Avoidance