paper-with-me

Papers

Audeo: Audio Generation for a Silent Performance Video

2020-06-23 · NeurIPS 2020 12 · Kun Su, Xiulong Liu, Eli Shlizerman

We present a novel system that gets as an input video frames of a musician playing the piano and generates the music for that video. Generation of music from visual cues is a challenging problem and it is not clear whether it is an attainable goal at all. Our main aim in this work is to explore the plausibility of such a transformation and to identify cues and components able to carry the association of sounds with visual events. To achieve the transformation we built a full pipeline named \textit{Audeo}' containing three components. We first translate the video frames of the keyboard and the musician hand movements into raw mechanical musical symbolic representation Piano-Roll (Roll) for each video frame which represents the keys pressed at each time step. We then adapt the Roll to be amenable for audio synthesis by including temporal correlations. This step turns out to be critical for meaningful audio generation. As a last step, we implement Midi synthesizers to generate realistic music. \textit{Audeo} converts video to audio smoothly and clearly with only a few setup constraints. We evaluate \textit{Audeo} on in the wild' piano performance videos and obtain that their generated music is of reasonable audio quality and can be successfully recognized with high precision by popular music identification software.

📄 PDF Abstract BibTeX arXiv:2006.14348

Code (1)

shlizee/Audeo 공식 구현 pytorch

Tasks

Audio GenerationAudio SynthesisVideo Generation

Similar Papers 제목 키워드 기반

MMAudio-LABEL: Audio Event Labeling via Audio Generation for Silent Video

2026-05-01 · Kazuya Tateishi, Akira Takahashi, Atsuo Hiroe, Hirofumi Takeda 외 arxiv

Recent advances in multimodal generation have enabled high-quality audio generation from silent videos. Practical applications, such as sound production, demand not only the generated audio but also explicit sound event …

Sound Event Detectionmultimodal generationAudio Generation

Synthesizing Audio from Silent Video using Sequence to Sequence Modeling

2024-04-25 · Hugo Garrido-Lestache Belinchon, Helina Mulugeta, Adam Haile

Generating audio from a video's visual context has multiple practical applications in improving how we interact with audio-visual media - for example, enhancing CCTV footage analysis, restoring historical videos (e.g., s…

DecoderDiversityVideo Generation

EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos

2024-07-30 · Aashish Rai, Srinath Sridhar

We introduce EgoSonics, a method to generate semantically meaningful and synchronized audio tracks conditioned on silent egocentric videos. Generating audio for silent egocentric videos could open new applications in vir…

Audio SynthesisVideo Summarization

ViSAGe: Video-to-Spatial Audio Generation

2025-06-13 · Jaeyeon Kim, Heeseung Yun, Gunhee Kim

Spatial audio is essential for enhancing the immersiveness of audio-visual experiences, yet its production typically demands complex recording systems and specialized expertise. In this work, we address a novel problem o…

Audio Generation

Assessing Identity Leakage in Talking Face Generation: Metrics and Evaluation Framework

2025-11-05 · Dogucan Yaman, Fevziye Irem Eyiokur, Hazım Kemal Ekenel, Alexander Waibel arxiv

Video editing-based talking face generation aims to preserve video details such as pose, lighting, and gestures while modifying only lip motion, often using an identity reference image to maintain speaker consistency. Ho…

Talking Face Generation