paper-with-me

Papers

Self-supervised Contrastive Video-Speech Representation Learning for Ultrasound

2020-08-14 · Jianbo Jiao, Yifan Cai, Mohammad Alsharid, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble

In medical imaging, manual annotations can be expensive to acquire and sometimes infeasible to access, making conventional deep learning-based models difficult to scale. As a result, it would be beneficial if useful representations could be derived from raw data without the need for manual annotations. In this paper, we propose to address the problem of self-supervised representation learning with multi-modal ultrasound video-speech raw data. For this case, we assume that there is a high correlation between the ultrasound video and the corresponding narrative speech audio of the sonographer. In order to learn meaningful representations, the model needs to identify such correlation and at the same time understand the underlying anatomical features. We designed a framework to model the correspondence between video and audio without any kind of human annotations. Within this framework, we introduce cross-modal contrastive learning and an affinity-aware self-paced learning scheme to enhance correlation modelling. Experimental evaluations on multi-modal fetal ultrasound video and audio show that the proposed approach is able to learn strong representations and transfers well to downstream tasks of standard plane detection and eye-gaze prediction.

📄 PDF Abstract BibTeX arXiv:2008.06607

Code (0)

등록된 구현이 없습니다.

Tasks

Contrastive LearningGaze PredictionRepresentation LearningSpeech Representation Learning

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음

Similar Papers 제목 키워드 기반

Speech SIMCLR: Combining Contrastive and Reconstruction Objective for Self-supervised Speech Representation Learning

2020-10-27 · Dongwei Jiang, Wubo Li, Miao Cao, Wei Zou 외

Self-supervised visual pretraining has shown significant progress recently. Among those methods, SimCLR greatly advanced the state of the art in self-supervised and semi-supervised learning on ImageNet. The input feature…

Emotion RecognitionRepresentation LearningSpeech Emotion Recognitionspeech-recognition+2

Multichannel AV-wav2vec2: A Framework for Learning Multichannel Multi-Modal Speech Representation

2024-01-07 · Qiushi Zhu, Jie Zhang, Yu Gu, Yuchen Hu 외

Self-supervised speech pre-training methods have developed rapidly in recent years, which show to be very effective for many near-field single-channel speech tasks. However, far-field multichannel speech processing is su…

Audio-Visual Speech RecognitionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Contrastive Learning+6

Learning Video Representations using Contrastive Bidirectional Transformer

2019-06-13 · Chen Sun, Fabien Baradel, Kevin Murphy, Cordelia Schmid

This paper proposes a self-supervised learning approach for video features that results in significantly improved performance on downstream tasks (such as video classification, captioning and segmentation) compared to ex…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation LearningSelf-Supervised Learning+4

Non-Contrastive Self-supervised Learning for Utterance-Level Information Extraction from Speech

2022-08-10 · Jaejin Cho, Jes'us Villalba, Laureano Moro-Velazquez, Najim Dehak

In recent studies, self-supervised pre-trained models tend to outperform supervised pre-trained models in transfer learning. In particular, self-supervised learning (SSL) of utterance-level speech representation can be u…

Alzheimer's Disease DetectionEmotion RecognitionSelf-Supervised LearningSpeaker Verification+2

Local plasticity rules can learn deep representations using self-supervised contrastive predictions

2020-10-16 · NeurIPS 2021 12 · Bernd Illing, Jean Ventura, Guillaume Bellec, Wulfram Gerstner

Learning in the brain is poorly understood and learning rules that respect biological constraints, yet yield deep hierarchical representations, are still unknown. Here, we propose a learning rule that takes inspiration f…