Papers Audio-Visual Synchronization
“Audio-Visual Synchronization” 태그가 달린 논문 32편 · 필터 해제
Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation
We propose Kling-Foley, a large-scale multimodal Video-to-Audio generation model that synthesizes high-quality audio synchronized with video content. In Kling-Foley, we introduce multimodal diffusion transformers to mode…
Audio GenerationAudio-Visual SynchronizationAudio-Sync Video Generation with Multi-Stream Temporal Control
Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translat…
Audio-Visual SynchronizationVideo AlignmentVideo GenerationOmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions
In this paper, we introduce Online Multimodal Conversational Response Generation (OMCRG), a novel task that aims to online generate synchronized verbal and non-verbal listener feedback, conditioned on the speaker's multi…
Audio-Visual SynchronizationConversational Response GenerationLarge Language ModelMultimodal Large Language Model+1CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization
The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tasks, especially in challenging conditions…
Active Speaker DetectionAudio-Visual Speech RecognitionAudio-Visual SynchronizationRepresentation Learning+4DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation
Currently, high-quality, synchronized audio is synthesized using various multi-modal joint learning frameworks, leveraging video and optional text inputs. In the video-to-audio benchmarks, video-to-audio quality, semanti…
Audio GenerationAudio-Visual Synchronizationtext-to-speechText to SpeechUniSync: A Unified Framework for Audio-Visual Synchronization
Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this challenge through rule-based approaches and…
Audio-Visual SynchronizationContrastive LearningFace GenerationFace Parsing+1FREAK: Frequency-modulated High-fidelity and Real-time Audio-driven Talking Portrait Synthesis
Achieving high-fidelity lip-speech synchronization in audio-driven talking portrait synthesis remains challenging. While multi-stage pipelines or diffusion models yield high-quality results, they suffer from high computa…
Audio-Visual SynchronizationMMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework MMAudio. In contrast to single-modality training conditioned on (l…
Audio GenerationAudio SynthesisAudio-Visual SynchronizationVideo-to-Sound GenerationMuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling
Real-time video dubbing that preserves identity consistency while achieving accurate lip synchronization remains a critical challenge. Existing approaches face a trilemma: diffusion-based methods achieve high visual fide…
Audio-Visual SynchronizationGPUVideo GenerationDraw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis
Foley is a term commonly used in filmmaking, referring to the addition of daily sound effects to silent films or videos to enhance the auditory experience. Video-to-Audio (V2A), as a particular type of automatic foley ta…
Audio SynthesisAudio-Visual SynchronizationA Comprehensive Review and Taxonomy of Audio-Visual Synchronization Techniques for Realistic Speech Animation
In many applications, synchronizing audio with visuals is crucial, such as in creating graphic animations for films or games, translating movie audio into different languages, and developing metaverse applications. This …
Audio-Visual SynchronizationRealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network
Person-generic audio-driven face generation is a challenging task in computer vision. Previous methods have achieved remarkable progress in audio-visual synchronization, but there is still a significant gap between curre…
Audio-Visual SynchronizationFace GenerationExplicit Correlation Learning for Generalizable Cross-Modal Deepfake Detection
With the rising prevalence of deepfakes, there is a growing interest in developing generalizable detection methods for various types of deepfakes. While effective in their specific modalities, traditional detection metho…
Audio-Visual SynchronizationDeepFake DetectionFace SwappingPEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores
Recent advancements in audio-visual generative modeling have been propelled by progress in deep learning and the availability of data-rich benchmarks. However, the growth is not attributed solely to models and benchmarks…
Audio-Visual SynchronizationSynchformer: Efficient Synchronization from Sparse Cues
Our objective is audio-visual synchronization with a focus on 'in-the-wild' videos, such as those on YouTube, where synchronization cues can be sparse. Our contributions include a novel audio-visual synchronization model…
Audio-Visual SynchronizationCoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing
There has been a long-standing quest for a unified audio-visual-text model to enable various multimodal understanding tasks, which mimics the listening, seeing and reading process of human beings. Humans tends to represe…
AudioCapsAudio-Visual SynchronizationLanguage ModelingLanguage Modelling+3Comparative Analysis of Deep-Fake Algorithms
Due to the widespread use of smartphones with high-quality digital cameras and easy access to a wide range of software apps for recording, editing, and sharing videos and images, as well as the deep learning AI platforms…
Audio-Visual SynchronizationDeepFake DetectionDeep LearningFace Swapping+1Audio-driven Talking Face Generation with Stabilized Synchronization Loss
Talking face generation aims to create realistic videos with accurate lip synchronization and high visual quality, using given audio and reference video while preserving identity and visual characteristics. In this paper…
Audio-Visual SynchronizationFace GenerationTalking Face GenerationTarget Active Speaker Detection with Audio-visual Cues
In active speaker detection (ASD), we would like to detect whether an on-screen person is speaking based on audio-visual cues. Previous studies have primarily focused on modeling audio-visual synchronization cue, which d…
Active Speaker DetectionAudio-Visual SynchronizationOn the Audio-visual Synchronization for Lip-to-Speech Synthesis
Most lip-to-speech (LTS) synthesis models are trained and evaluated under the assumption that the audio-video pairs in the dataset are perfectly synchronized. In this work, we show that the commonly used audio-visual dat…
Audio-Visual SynchronizationLip to Speech SynthesisSpeech Synthesis