LAV: Audio-Driven Dynamic Visual Generation with Neural Compression and StyleGAN2
This paper introduces LAV (Latent Audio-Visual), a system that integrates EnCodec's neural audio compression with StyleGAN2's generative capabilities to produce visually dynamic outputs driven by pre-recorded audio. Unlike previous works that rely on explicit feature mappings, LAV uses EnCodec embeddings as latent representations, directly transformed into StyleGAN2's style latent space via randomly initialized linear mapping. This approach preserves semantic richness in the transformation, enabling nuanced and semantically coherent audio-visual translations. The framework demonstrates the potential of using pretrained audio compression models for artistic and computational applications.
Code (0)
등록된 구현이 없습니다.
Tasks
Audio CompressionSimilar Papers 제목 키워드 기반
DASH: Dynamic Audio-Driven Semantic Chunking for Efficient Omnimodal Token Compression
Omnimodal large language models (OmniLLMs) jointly process audio and visual streams, but the resulting long multimodal token sequences make inference prohibitively expensive. Existing compression methods typically rely o…
Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos
Talking head video compression has advanced with neural rendering and keypoint-based methods, but challenges remain, especially at low bit rates, including handling large head movements, suboptimal lip synchronization, a…
Neural RenderingVideo CompressionDual Audio-Centric Modality Coupling for Talking Head Generation
The generation of audio-driven talking head videos is a key challenge in computer vision and graphics, with applications in virtual avatars and digital media. Traditional approaches often struggle with capturing the comp…
NeRFTalking Head Generationtext-to-speechText to SpeechAudio-Visual Target Speaker Enhancement on Multi-Talker Environment using Event-Driven Cameras
We propose a method to address audio-visual target speaker enhancement in multi-talker environments using event-driven cameras. State of the art audio-visual speech separation methods shows that crucial information is th…
Optical Flow EstimationSpeech SeparationDo Joint Audio-Video Generation Models Understand Physics?
Joint audio-video generation models are rapidly approaching professional production quality, raising a central question: do they understand audio-visual physics, or merely generate plausible sounds and frames that violat…
Video Generation