paper-with-me

Papers

EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos

2024-07-30 · Aashish Rai, Srinath Sridhar

We introduce EgoSonics, a method to generate semantically meaningful and synchronized audio tracks conditioned on silent egocentric videos. Generating audio for silent egocentric videos could open new applications in virtual reality, assistive technologies, or for augmenting existing datasets. Existing work has been limited to domains like speech, music, or impact sounds and cannot capture the broad range of audio frequencies found in egocentric videos. EgoSonics addresses these limitations by building on the strengths of latent diffusion models for conditioned audio synthesis. We first encode and process paired audio-video data to make them suitable for generation. The encoded data is then used to train a model that can generate an audio track that captures the semantics of the input video. Our proposed SyncroNet builds on top of ControlNet to provide control signals that enables generation of temporally synchronized audio. Extensive evaluations and a comprehensive user study show that our model outperforms existing work in audio quality, and in our proposed synchronization evaluation method. Furthermore, we demonstrate downstream applications of our model in improving video summarization.

📄 PDF Abstract BibTeX arXiv:2407.20592

Code (0)

등록된 구현이 없습니다.

Tasks

Audio SynthesisVideo Summarization

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Diff-Foley: Synchronized Video-to-Audio Synthesis with Latent Diffusion Models

2023-06-29 · NeurIPS 2023 11 · Simian Luo, Chuanhao Yan, Chenxu Hu, Hang Zhao

The Video-to-Audio (V2A) model has recently gained attention for its practical application in generating audio directly from silent videos, particularly in video/film production. However, previous methods in V2A have lim…

Audio Synthesis

Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos

2024-06-13 · Changan Chen, Puyuan Peng, Ami Baid, Zihui Xue 외

Generating realistic audio for human actions is important for many applications, such as creating sound effects for films or virtual reality games. Existing approaches implicitly assume total correspondence between the v…

Audio GenerationRetrieval-augmented Generation

SyncAnimation: A Real-Time End-to-End Framework for Audio-Driven Human Pose and Talking Head Animation

2025-01-24 · Yujian Liu, Shidang Xu, Jing Guo, Dingbin Wang 외

Generating talking avatar driven by audio remains a significant challenge. Existing methods typically require high computational costs and often lack sufficient facial detail and realism, making them unsuitable for appli…

NeRF

LTX-2: Efficient Joint Audio-Visual Foundation Model

2026-01-06 · Yoav HaCohen, Benny Brazowski, Nisan Chiprut, Yaki Bitterman 외 arxiv

Recent text-to-video diffusion models can generate compelling video sequences, yet they remain silent -- missing the semantic, emotional, and atmospheric cues that audio provides. We introduce LTX-2, an open-source found…

Audio GenerationVideo Generation

SALSA-V: Shortcut-Augmented Long-form Synchronized Audio from Videos

2025-10-03 · Amir Dellali, Luca A. Lanzendörfer, Florian Grötschla, Roger Wattenhofer arxiv

We propose SALSA-V, a multimodal video-to-audio generation model capable of synthesizing highly synchronized, high-fidelity long-form audio from silent video content. Our approach introduces a masked diffusion objective,…

Audio Generation