Papers Video Synchronization
“Video Synchronization” 태그가 달린 논문 30편 · 필터 해제
Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization
Video synchronization-aligning multiple video streams capturing the same event from different angles-is crucial for applications such as reality TV show production, sports analysis, surveillance, and autonomous systems. …
Pose EstimationVideo SynchronizationAlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation
In this paper, we address the task of multimodal-to-speech generation, which aims to synthesize high-quality speech from multiple input modalities: text, video, and reference audio. This task has gained increasing attent…
In-Context LearningSpeech SynthesisVideo SynchronizationVoice SimilarityDRIFT open dataset: A drone-derived intelligence for traffic analysis in urban environmen
Reliable traffic data are essential for understanding urban mobility and developing effective traffic management strategies. This study introduces the DRone-derived Intelligence For Traffic analysis (DRIFT) dataset, a la…
object-detectionObject DetectionVideo SynchronizationMultimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models
Advances in generative artificial intelligence have altered multimedia creation, allowing for automatic cinematic video synthesis from text inputs. This work describes a method for creating 60-second cinematic movies inc…
Audio GenerationGPUImage GenerationVideo SynchronizationOmniTalker: Real-Time Text-Driven Talking Head Generation with In-Context Audio-Visual Style Replication
Recent years have witnessed remarkable advances in talking head generation, owing to its potential to revolutionize the human-AI interaction from text interfaces into realistic video chats. However, research on text-driv…
Talking Head GenerationVideo SynchronizationAV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation
We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioni…
Video GenerationVideo SynchronizationFoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
We study Neural Foley, the automatic generation of high-quality sound effects synchronizing with videos, enabling an immersive audio-visual experience. Despite its wide range of applications, existing approaches encounte…
Audio GenerationVideo AlignmentVideo SynchronizationCollaborative Video Diffusion: Consistent Multi-video Generation with Camera Control
Research on video generation has recently made tremendous progress, enabling high-quality videos to be generated from text prompts or images. Adding control to the video generation process is an important goal moving for…
Scene GenerationVideo GenerationVideo SynchronizationDance Any Beat: Blending Beats with Visuals in Dance Video Generation
Generating dance from music is crucial for advancing automated choreography. Current methods typically produce skeleton keypoint sequences instead of dance videos and lack the capability to make specific individuals danc…
Image to Video GenerationOptical Flow EstimationRhythmVideo Generation+1Context-aware Talking Face Video Generation
In this paper, we consider a novel and practical case for talking face video generation. Specifically, we focus on the scenarios involving multi-people interactions, where the talking context, such as audience or surroun…
Video GenerationVideo SynchronizationPoseSync: Robust pose based video synchronization
Pose based video sychronization can have applications in multiple domains such as gameplay performance evaluation, choreography or guiding athletes. The subject's actions could be compared and evaluated against those per…
Dynamic Time WarpingVideo SynchronizationVideo alignment using unsupervised learning of local and global features
In this paper, we tackle the problem of video alignment, the process of matching the frames of a pair of videos containing similar actions. The main challenge in video alignment is that accurate correspondence should be …
Dynamic Time WarpingHuman DetectionPose EstimationTime Series+2Deep learning-based stereo camera multi-video synchronization
Stereo vision is essential for many applications. Currently, the synchronization of the streams coming from two cameras is done using mostly hardware. A software-based synchronization method would reduce the cost, weight…
Deep LearningVideo SynchronizationModEFormer: Modality-Preserving Embedding for Audio-Video Synchronization using Transformers
Lack of audio-video synchronization is a common problem during television broadcasts and video conferencing, leading to an unsatisfactory viewing experience. A widely accepted paradigm is to create an error detection mec…
Contrastive LearningVideo SynchronizationBronchoscopic video synchronization for interactive multimodal inspection of bronchial lesions
With lung cancer being the most fatal cancer worldwide, it is important to detect the disease early. A potentially effective way of detecting early cancer lesions developing along the airway walls (epithelium) is broncho…
Computed Tomography (CT)Video SynchronizationApplying Automated Machine Translation to Educational Video Courses
We studied the capability of automated machine translation in the online video education space by automatically translating Khan Academy videos with state-of-the-art translation models and applying text-to-speech synthes…
Machine TranslationSpeech Synthesistext-to-speechText to Speech+3ReVISE: Self-Supervised Speech Resynthesis With Visual Input for Universal and Generalized Speech Regeneration
Prior works on improving speech quality with visual input typically study each type of auditory distortion separately (e.g., separation, inpainting, video-to-speech) and present tailored algorithms. This paper propos…
Audio-Visual Speech RecognitionResynthesisspeech-recognitionSpeech Recognition+6SIDGAN: High-Resolution Dubbed Video Generation via Shift-Invariant Learning
Dubbed video generation aims to accurately synchronize mouth movements of a given facial video with driving audio while preserving identity and scene-specific visual dynamics, such as head pose and lighting. Despite …
Image GenerationVideo GenerationVideo SynchronizationReVISE: Self-Supervised Speech Resynthesis with Visual Input for Universal and Generalized Speech Enhancement
Prior works on improving speech quality with visual input typically study each type of auditory distortion separately (e.g., separation, inpainting, video-to-speech) and present tailored algorithms. This paper proposes t…
Audio-Visual Speech RecognitionResynthesisSpeech Enhancementspeech-recognition+7A subjective study of the perceptual acceptability of audio-video desynchronization in sports videos
This paper presents the results of a study conducted on the perceptual acceptability of audio-video desynchronization for sports videos. The study was conducted with 45 videos generated by applying 8 audio-video offsets …
Video Synchronization