paper-with-me

홈 › Papers

Strumming to the Beat: Audio-Conditioned Contrastive Video Textures

2021-04-06 · Medhini Narasimhan, Shiry Ginosar, Andrew Owens, Alexei A. Efros, Trevor Darrell

We introduce a non-parametric approach for infinite video texture synthesis using a representation learned via contrastive learning. We take inspiration from Video Textures, which showed that plausible new videos could be generated from a single one by stitching its frames together in a novel yet consistent order. This classic work, however, was constrained by its use of hand-designed distance metrics, limiting its use to simple, repetitive videos. We draw on recent techniques from self-supervised learning to learn this distance metric, allowing us to compare frames in a manner that scales to more challenging dynamics, and to condition on other data, such as audio. We learn representations for video frames and frame-to-frame transition probabilities by fitting a video-specific model trained using contrastive learning. To synthesize a texture, we randomly sample frames with high transition probabilities to generate diverse temporally smooth videos with novel sequences and transitions. The model naturally extends to an audio-conditioned setting without requiring any finetuning. Our model outperforms baselines on human perceptual scores, can handle a diverse range of input videos, and can combine semantic and audio-visual cues in order to synthesize videos that synchronize well with an audio signal.

📄 PDF Abstract BibTeX arXiv:2104.02687

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningSelf-Supervised LearningTexture Synthesis

Similar Papers 제목 키워드 기반

Joint Transcription of Acoustic Guitar Strumming Directions and Chords

2025-08-11 · Sebastian Murgul, Johannes Schimper, Michael Heizmann arxiv

Automatic transcription of guitar strumming is an underrepresented and challenging task in Music Information Retrieval (MIR), particularly for extracting both strumming directions and chord progressions from audio signal…

Information RetrievalAction Detection

Contrastive Video Textures

2021-01-01 · Medhini Narasimhan, Shiry Ginosar, Andrew Owens, Alexei A Efros 외

Existing methods for video generation struggle to generate more than a short sequence of frames. We introduce a non-parametric approach for infinite video generation based on learning to resample frames from an input vid…

Contrastive LearningVideo Generation

DiffGAP: A Lightweight Diffusion Module in Contrastive Space for Bridging Cross-Model Gap

2025-03-15 · Shentong Mo, Zehua Chen, Fan Bao, Jun Zhu

Recent works in cross-modal understanding and generation, notably through models like CLAP (Contrastive Language-Audio Pretraining) and CAVP (Contrastive Audio-Visual Pretraining), have significantly enhanced the alignme…

AudioCapsAudio GenerationDenoising

MotionBeat: Motion-Aligned Music Representation via Embodied Contrastive Learning and Bar-Equivariant Contact-Aware Encoding

2025-10-15 · Xuanchen Wang, Heng Wang, Weidong Cai arxiv

Music is both an auditory and an embodied phenomenon, closely linked to human motion and naturally expressed through dance. However, most existing audio representations neglect this embodied dimension, limiting their abi…

Representation LearningContrastive LearningEmotion RecognitionBeat Tracking

Language-Based Audio Retrieval with Converging Tied Layers and Contrastive Loss

2022-06-29 · Andrew Koh, Eng Siong Chng

In this paper, we tackle the new Language-Based Audio Retrieval task proposed in DCASE 2022. Firstly, we introduce a simple, scalable architecture which ties both the audio and text encoder together. Secondly, we show th…

Retrieval