SyncNet: Using Causal Convolutions and Correlating Objective for Time Delay Estimation in Audio Signals
This paper addresses the task of performing robust and reliable time-delay estimation in audio-signals in noisy and reverberating environments. In contrast to the popular signal processing based methods, this paper proposes machine learning based method, i.e., a semi-causal convolutional neural network consisting of a set of causal and anti-causal layers with a novel correlation-based objective function. The causality in the network ensures non-leakage of representations from future time-intervals and the proposed loss function makes the network generate sequences with high correlation at the actual time delay. The proposed approach is also intrinsically interpretable as it does not lose time information. Even a shallow convolution network is able to capture local patterns in sequences, while also correlating them globally. SyncNet outperforms other classical approaches in estimating mutual time delays for different types of audio signals including pulse, speech and musical beats.
Code (1)
Methods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Lost in Time? A Meta-Learning Framework for Time-Shift-Tolerant Physiological Signal Transformation
Translating non-invasive signals such as photoplethysmography (PPG) and ballistocardiography (BCG) into clinically meaningful signals like arterial blood pressure (ABP) is vital for continuous, low-cost healthcare monito…
Learning with noisy labelsEcho-SyncNet: Self-supervised Cardiac View Synchronization in Echocardiography
In echocardiography (echo), an electrocardiogram (ECG) is conventionally used to temporally align different cardiac views for assessing critical measurements. However, in emergencies or point-of-care situations, acquirin…
One-Shot LearningSelf-Supervised LearningInfoSyncNet: Information Synchronization Temporal Convolutional Network for Visual Speech Recognition
Estimating spoken content from silent videos is crucial for applications in Assistive Technology (AT) and Augmented Reality (AR). However, accurately mapping lip movement sequences in videos to words poses significant ch…
Visual Speech RecognitionData AugmentationLatentSync: Audio Conditioned Latent Diffusion Models for Lip Sync
We present LatentSync, an end-to-end lip sync framework based on audio conditioned latent diffusion models without any intermediate motion representation, diverging from previous diffusion-based lip sync methods based on…
Portrait AnimationAudio-driven Talking Face Generation with Stabilized Synchronization Loss
Talking face generation aims to create realistic videos with accurate lip synchronization and high visual quality, using given audio and reference video while preserving identity and visual characteristics. In this paper…
Audio-Visual SynchronizationFace GenerationTalking Face Generation