paper-with-me

홈 › Papers

Synchronising audio and ultrasound by learning cross-modal embeddings

2019-07-01 · Aciel Eshky, Manuel Sam Ribeiro, Korin Richmond, Steve Renals

Audiovisual synchronisation is the task of determining the time offset between speech audio and a video recording of the articulators. In child speech therapy, audio and ultrasound videos of the tongue are captured using instruments which rely on hardware to synchronise the two modalities at recording time. Hardware synchronisation can fail in practice, and no mechanism exists to synchronise the signals post hoc. To address this problem, we employ a two-stream neural network which exploits the correlation between the two modalities to find the offset. We train our model on recordings from 69 speakers, and show that it correctly synchronises 82.9% of test utterances from unseen therapy sessions and unseen speakers, thus considerably reducing the number of utterances to be manually synchronised. An analysis of model performance on the test utterances shows that directed phone articulations are more difficult to automatically synchronise compared to utterances containing natural variation in speech such as words, sentences, or conversations.

📄 PDF Abstract BibTeX arXiv:1907.00758

Code (1)

aeshky/ultrasync 공식 구현

Similar Papers 제목 키워드 기반

Automatic audiovisual synchronisation for ultrasound tongue imaging

2021-05-31 · Aciel Eshky, Joanne Cleland, Manuel Sam Ribeiro, Eleanor Sugden 외

Ultrasound tongue imaging is used to visualise the intra-oral articulators during speech production. It is utilised in a range of applications, including speech and language therapy and phonetics research. Ultrasound and…

Self-supervised Contrastive Video-Speech Representation Learning for Ultrasound

2020-08-14 · Jianbo Jiao, Yifan Cai, Mohammad Alsharid, Lior Drukker 외

In medical imaging, manual annotations can be expensive to acquire and sometimes infeasible to access, making conventional deep learning-based models difficult to scale. As a result, it would be beneficial if useful repr…

Contrastive LearningGaze PredictionRepresentation LearningSpeech Representation Learning

Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors

2022-10-13 · Vladimir Iashin, Weidi Xie, Esa Rahtu, Andrew Zisserman

The objective of this paper is audio-visual synchronisation of general videos 'in the wild'. For such videos, the events that may be harnessed for synchronisation cues may be spatially small and may occur only infrequent…

Audio-Visual Synchronization

Steganography Beyond Space-Time with Chain of Multimodal AI

2025-02-25 · Ching-Chun Chang, Isao Echizen

Steganography is the art and science of covert writing, with a broad range of applications interwoven within the realm of cybersecurity. As artificial intelligence continues to evolve, its ability to synthesise realistic…

Face SwappingText GenerationVoice Cloning

AVGZSLNet: Audio-Visual Generalized Zero-Shot Learning by Reconstructing Label Features from Multi-Modal Embeddings

2020-05-27 · Pratik Mazumder, Pravendra Singh, Kranti Kumar Parida, Vinay P. Namboodiri

In this paper, we propose a novel approach for generalized zero-shot learning in a multi-modal setting, where we have novel classes of audio/video during testing that are not seen during training. We use the semantic rel…

DecoderGeneralized Zero-Shot LearningGZSL Video ClassificationRetrieval+3