paper-with-me

Papers Audio-Visual Synchronization

“Audio-Visual Synchronization” 태그가 달린 논문 32편 · 필터 해제

Kling-Foley: Multimodal Diffusion Transformer for High-Quality Video-to-Audio Generation

2025-06-24 · Jun Wang, Xijuan Zeng, Chunyu Qiang, Ruilong Chen 외

We propose Kling-Foley, a large-scale multimodal Video-to-Audio generation model that synthesizes high-quality audio synchronized with video content. In Kling-Foley, we introduce multimodal diffusion transformers to mode…

Audio GenerationAudio-Visual Synchronization

Audio-Sync Video Generation with Multi-Stream Temporal Control

2025-06-09 · Shuchen Weng, Haojie Zheng, Zheng Chang, Si Li 외

Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translat…

Audio-Visual SynchronizationVideo AlignmentVideo Generation

OmniResponse: Online Multimodal Conversational Response Generation in Dyadic Interactions

2025-05-27 · Cheng Luo, Jianghui Wang, Bing Li, Siyang Song 외

In this paper, we introduce Online Multimodal Conversational Response Generation (OMCRG), a novel task that aims to online generate synchronized verbal and non-verbal listener feedback, conditioned on the speaker's multi…

Audio-Visual SynchronizationConversational Response GenerationLarge Language ModelMultimodal Large Language Model+1

CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization

2025-05-06 · Detao Bai, Zhiheng Ma, Xihan Wei, Liefeng Bo

The inherent synchronization between a speaker's lip movements, voice, and the underlying linguistic content offers a rich source of information for improving speech processing tasks, especially in challenging conditions…

Active Speaker DetectionAudio-Visual Speech RecognitionAudio-Visual SynchronizationRepresentation Learning+4

DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation

2025-03-28 · Haomin Zhang, Chang Liu, Junjie Zheng, Zihao Chen 외

Currently, high-quality, synchronized audio is synthesized using various multi-modal joint learning frameworks, leveraging video and optional text inputs. In the video-to-audio benchmarks, video-to-audio quality, semanti…

Audio GenerationAudio-Visual Synchronizationtext-to-speechText to Speech

UniSync: A Unified Framework for Audio-Visual Synchronization

2025-03-20 · Tao Feng, Yifan Xie, Xun Guan, Jiyuan Song 외

Precise audio-visual synchronization in speech videos is crucial for content quality and viewer comprehension. Existing methods have made significant strides in addressing this challenge through rule-based approaches and…

Audio-Visual SynchronizationContrastive LearningFace GenerationFace Parsing+1

FREAK: Frequency-modulated High-fidelity and Real-time Audio-driven Talking Portrait Synthesis

2025-03-06 · Ziqi Ni, Ao Fu, Yi Zhou

Achieving high-fidelity lip-speech synchronization in audio-driven talking portrait synthesis remains challenging. While multi-stage pipelines or diffusion models yield high-quality results, they suffer from high computa…

Audio-Visual Synchronization

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

2024-12-19 · CVPR 2025 1 · Ho Kei Cheng, Masato Ishii, Akio Hayakawa, Takashi Shibuya 외

We propose to synthesize high-quality and synchronized audio, given video and optional text conditions, using a novel multimodal joint training framework MMAudio. In contrast to single-modality training conditioned on (l…

Audio GenerationAudio SynthesisAudio-Visual SynchronizationVideo-to-Sound Generation

MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

2024-10-14 · Yue Zhang, Zhizhou Zhong, Minhao Liu, Zhaokang Chen 외

Real-time video dubbing that preserves identity consistency while achieving accurate lip synchronization remains a critical challenge. Existing approaches face a trilemma: diffusion-based methods achieve high visual fide…

Audio-Visual SynchronizationGPUVideo Generation

Draw an Audio: Leveraging Multi-Instruction for Video-to-Audio Synthesis

2024-09-10 · Qi Yang, Binjie Mao, Zili Wang, Xing Nie 외

Foley is a term commonly used in filmmaking, referring to the addition of daily sound effects to silent films or videos to enhance the auditory experience. Video-to-Audio (V2A), as a particular type of automatic foley ta…

Audio SynthesisAudio-Visual Synchronization

A Comprehensive Review and Taxonomy of Audio-Visual Synchronization Techniques for Realistic Speech Animation

2024-07-24 · Jose Geraldo Fernandes, Sinval Nascimento, Daniel Dominguete, André Oliveira 외

In many applications, synchronizing audio with visuals is crucial, such as in creating graphic animations for films or games, translating movie audio into different languages, and developing metaverse applications. This …

Audio-Visual Synchronization

RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network

2024-06-26 · Xiaozhong Ji, Chuming Lin, Zhonggan Ding, Ying Tai 외

Person-generic audio-driven face generation is a challenging task in computer vision. Previous methods have achieved remarkable progress in audio-visual synchronization, but there is still a significant gap between curre…

Audio-Visual SynchronizationFace Generation

Explicit Correlation Learning for Generalizable Cross-Modal Deepfake Detection

2024-04-30 · Cai Yu, Shan Jia, Xiaomeng Fu, Jin Liu 외

With the rising prevalence of deepfakes, there is a growing interest in developing generalizable detection methods for various types of deepfakes. While effective in their specific modalities, traditional detection metho…

Audio-Visual SynchronizationDeepFake DetectionFace Swapping

PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores

2024-04-10 · Lucas Goncalves, Prashant Mathur, Chandrashekhar Lavania, Metehan Cekic 외

Recent advancements in audio-visual generative modeling have been propelled by progress in deep learning and the availability of data-rich benchmarks. However, the growth is not attributed solely to models and benchmarks…

Audio-Visual Synchronization

Synchformer: Efficient Synchronization from Sparse Cues

2024-01-29 · Vladimir Iashin, Weidi Xie, Esa Rahtu, Andrew Zisserman

Our objective is audio-visual synchronization with a focus on 'in-the-wild' videos, such as those on YouTube, where synchronization cues can be sparse. Our contributions include a novel audio-visual synchronization model…

Audio-Visual Synchronization

CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing

2024-01-22 · Xianghu Yue, Xiaohai Tian, Lu Lu, Malu Zhang 외

There has been a long-standing quest for a unified audio-visual-text model to enable various multimodal understanding tasks, which mimics the listening, seeing and reading process of human beings. Humans tends to represe…

AudioCapsAudio-Visual SynchronizationLanguage ModelingLanguage Modelling+3

Comparative Analysis of Deep-Fake Algorithms

2023-09-06 · Nikhil Sontakke, Sejal Utekar, Shivansh Rastogi, Shriraj Sonawane

Due to the widespread use of smartphones with high-quality digital cameras and easy access to a wide range of software apps for recording, editing, and sharing videos and images, as well as the deep learning AI platforms…

Audio-Visual SynchronizationDeepFake DetectionDeep LearningFace Swapping+1

Audio-driven Talking Face Generation with Stabilized Synchronization Loss

2023-07-18 · Dogucan Yaman, Fevziye Irem Eyiokur, Leonard Bärmann, Hazim Kemal Ekenel 외

Talking face generation aims to create realistic videos with accurate lip synchronization and high visual quality, using given audio and reference video while preserving identity and visual characteristics. In this paper…

Audio-Visual SynchronizationFace GenerationTalking Face Generation

Target Active Speaker Detection with Audio-visual Cues

2023-05-22 · Yidi Jiang, Ruijie Tao, Zexu Pan, Haizhou Li

In active speaker detection (ASD), we would like to detect whether an on-screen person is speaking based on audio-visual cues. Previous studies have primarily focused on modeling audio-visual synchronization cue, which d…

Active Speaker DetectionAudio-Visual Synchronization

On the Audio-visual Synchronization for Lip-to-Speech Synthesis

2023-03-01 · ICCV 2023 1 · Zhe Niu, Brian Mak

Most lip-to-speech (LTS) synthesis models are trained and evaluated under the assumption that the audio-video pairs in the dataset are perfectly synchronized. In this work, we show that the commonly used audio-visual dat…

Audio-Visual SynchronizationLip to Speech SynthesisSpeech Synthesis
1–20 / 32 다음 →