paper-with-me

Papers Video Synchronization

“Video Synchronization” 태그가 달린 논문 30편 · 필터 해제

Beyond Audio and Pose: A General-Purpose Framework for Video Synchronization

2025-06-19 · Yosub Shin, Igor Molybog

Video synchronization-aligning multiple video streams capturing the same event from different angles-is crucial for applications such as reality TV show production, sports analysis, surveillance, and autonomous systems. …

Pose EstimationVideo Synchronization

AlignDiT: Multimodal Aligned Diffusion Transformer for Synchronized Speech Generation

2025-04-29 · Jeongsoo Choi, Ji-Hoon Kim, Kim Sung-Bin, Tae-Hyun Oh 외

In this paper, we address the task of multimodal-to-speech generation, which aims to synthesize high-quality speech from multiple input modalities: text, video, and reference audio. This task has gained increasing attent…

In-Context LearningSpeech SynthesisVideo SynchronizationVoice Similarity

DRIFT open dataset: A drone-derived intelligence for traffic analysis in urban environmen

2025-04-15 · Hyejin Lee, Seokjun Hong, Jeonghoon Song, Haechan Cho 외

Reliable traffic data are essential for understanding urban mobility and developing effective traffic management strategies. This study introduces the DRone-derived Intelligence For Traffic analysis (DRIFT) dataset, a la…

object-detectionObject DetectionVideo Synchronization

Multimodal Cinematic Video Synthesis Using Text-to-Image and Audio Generation Models

2025-04-06 · Sridhar S, Nithin A, Shakeel Rifath, Vasantha Raj

Advances in generative artificial intelligence have altered multimedia creation, allowing for automatic cinematic video synthesis from text inputs. This work describes a method for creating 60-second cinematic movies inc…

Audio GenerationGPUImage GenerationVideo Synchronization

OmniTalker: Real-Time Text-Driven Talking Head Generation with In-Context Audio-Visual Style Replication

2025-04-03 · Zhongjian Wang, Peng Zhang, Jinwei Qi, Guangyuan Wang Sheng Xu 외

Recent years have witnessed remarkable advances in talking head generation, owing to its potential to revolutionize the human-AI interaction from text interfaces into realistic video chats. However, research on text-driv…

Talking Head GenerationVideo Synchronization

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation

2024-12-19 · Moayed Haji-Ali, Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov 외

We propose AV-Link, a unified framework for Video-to-Audio (A2V) and Audio-to-Video (A2V) generation that leverages the activations of frozen video and audio diffusion models for temporally-aligned cross-modal conditioni…

Video GenerationVideo Synchronization

FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

2024-07-01 · Yiming Zhang, Yicheng Gu, Yanhong Zeng, Zhening Xing 외

We study Neural Foley, the automatic generation of high-quality sound effects synchronizing with videos, enabling an immersive audio-visual experience. Despite its wide range of applications, existing approaches encounte…

Audio GenerationVideo AlignmentVideo Synchronization

Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control

2024-05-27 · Zhengfei Kuang, Shengqu Cai, Hao He, Yinghao Xu 외

Research on video generation has recently made tremendous progress, enabling high-quality videos to be generated from text prompts or images. Adding control to the video generation process is an important goal moving for…

Scene GenerationVideo GenerationVideo Synchronization

Dance Any Beat: Blending Beats with Visuals in Dance Video Generation

2024-05-15 · Xuanchen Wang, Heng Wang, Dongnan Liu, Weidong Cai

Generating dance from music is crucial for advancing automated choreography. Current methods typically produce skeleton keypoint sequences instead of dance videos and lack the capability to make specific individuals danc…

Image to Video GenerationOptical Flow EstimationRhythmVideo Generation+1

Context-aware Talking Face Video Generation

2024-02-28 · Meidai Xuanyuan, Yuwang Wang, Honglei Guo, Qionghai Dai

In this paper, we consider a novel and practical case for talking face video generation. Specifically, we focus on the scenarios involving multi-people interactions, where the talking context, such as audience or surroun…

Video GenerationVideo Synchronization

PoseSync: Robust pose based video synchronization

2023-08-24 · Rishit Javia, Falak Shah, Shivam Dave

Pose based video sychronization can have applications in multiple domains such as gameplay performance evaluation, choreography or guiding athletes. The subject's actions could be compared and evaluated against those per…

Dynamic Time WarpingVideo Synchronization

Video alignment using unsupervised learning of local and global features

2023-04-13 · Niloufar Fakhfour, Mohammad ShahverdiKondori, Sajjad Hashembeiki, Mohammadjavad Norouzi 외

In this paper, we tackle the problem of video alignment, the process of matching the frames of a pair of videos containing similar actions. The main challenge in video alignment is that accurate correspondence should be …

Dynamic Time WarpingHuman DetectionPose EstimationTime Series+2

Deep learning-based stereo camera multi-video synchronization

2023-03-22 · Nicolas Boizard, Kevin El Haddad, Thierry Ravet, François Cresson 외

Stereo vision is essential for many applications. Currently, the synchronization of the streams coming from two cameras is done using mostly hardware. A software-based synchronization method would reduce the cost, weight…

Deep LearningVideo Synchronization

ModEFormer: Modality-Preserving Embedding for Audio-Video Synchronization using Transformers

2023-03-21 · Akash Gupta, Rohun Tripathi, WonDong Jang

Lack of audio-video synchronization is a common problem during television broadcasts and video conferencing, leading to an unsatisfactory viewing experience. A widely accepted paradigm is to create an error detection mec…

Contrastive LearningVideo Synchronization

Bronchoscopic video synchronization for interactive multimodal inspection of bronchial lesions

2023-03-20 · Qi Chang, Patrick D. Byrnes, Danish Ahmad, Jennifer Toth 외

With lung cancer being the most fatal cancer worldwide, it is important to detect the disease early. A potentially effective way of detecting early cancer lesions developing along the airway walls (epithelium) is broncho…

Computed Tomography (CT)Video Synchronization

Applying Automated Machine Translation to Educational Video Courses

2023-01-09 · Linden Wang

We studied the capability of automated machine translation in the online video education space by automatically translating Khan Academy videos with state-of-the-art translation models and applying text-to-speech synthes…

Machine TranslationSpeech Synthesistext-to-speechText to Speech+3

ReVISE: Self-Supervised Speech Resynthesis With Visual Input for Universal and Generalized Speech Regeneration

2023-01-01 · CVPR 2023 1 · Wei-Ning Hsu, Tal Remez, Bowen Shi, Jacob Donley 외

Prior works on improving speech quality with visual input typically study each type of auditory distortion separately (e.g., separation, inpainting, video-to-speech) and present tailored algorithms. This paper propos…

Audio-Visual Speech RecognitionResynthesisspeech-recognitionSpeech Recognition+6

SIDGAN: High-Resolution Dubbed Video Generation via Shift-Invariant Learning

2023-01-01 · ICCV 2023 1 · Urwa Muaz, WonDong Jang, Rohun Tripathi, Santhosh Mani 외

Dubbed video generation aims to accurately synchronize mouth movements of a given facial video with driving audio while preserving identity and scene-specific visual dynamics, such as head pose and lighting. Despite …

Image GenerationVideo GenerationVideo Synchronization

ReVISE: Self-Supervised Speech Resynthesis with Visual Input for Universal and Generalized Speech Enhancement

2022-12-21 · Wei-Ning Hsu, Tal Remez, Bowen Shi, Jacob Donley 외

Prior works on improving speech quality with visual input typically study each type of auditory distortion separately (e.g., separation, inpainting, video-to-speech) and present tailored algorithms. This paper proposes t…

Audio-Visual Speech RecognitionResynthesisSpeech Enhancementspeech-recognition+7

A subjective study of the perceptual acceptability of audio-video desynchronization in sports videos

2022-12-03 · Joshua Peter Ebenezer

This paper presents the results of a study conducted on the perceptual acceptability of audio-video desynchronization for sports videos. The study was conducted with 45 videos generated by applying 8 audio-video offsets …

Video Synchronization
1–20 / 30 다음 →