Self-supervised reinforcement learning for speaker localisation with the iCub humanoid robot
In the future robots will interact more and more with humans and will have to communicate naturally and efficiently. Automatic speech recognition systems (ASR) will play an important role in creating natural interactions and making robots better companions. Humans excel in speech recognition in noisy environments and are able to filter out noise. Looking at a person's face is one of the mechanisms that humans rely on when it comes to filtering speech in such noisy environments. Having a robot that can look toward a speaker could benefit ASR performance in challenging environments. To this aims, we propose a self-supervised reinforcement learning-based framework inspired by the early development of humans to allow the robot to autonomously create a dataset that is later used to learn to localize speakers with a deep learning network.
Code (0)
등록된 구현이 없습니다.
Tasks
Automatic Speech RecognitionAutomatic Speech Recognition (ASR)reinforcement-learningReinforcement Learning (RL)speech-recognitionSpeech RecognitionSimilar Papers 제목 키워드 기반
Spatial HuBERT: Self-supervised Spatial Speech Representation Learning for a Single Talker from Multi-channel Audio
Self-supervised learning has been used to leverage unlabelled data, improving accuracy and generalisation of speech systems through the training of representation models. While many recent works have sought to produce ef…
Representation LearningSelf-Supervised LearningSpeech Representation LearningWeakly supervised localisation of prostate cancer using reinforcement learning for bi-parametric MR images
In this paper we propose a reinforcement learning based weakly supervised system for localisation. We train a controller function to localise regions of interest within an image by introducing a novel reward definition t…
Multiple Instance LearningObjectSPARS: Self-Play Adversarial Reinforcement Learning for Segmentation of Liver Tumours
Accurate tumour segmentation is vital for various targeted diagnostic and therapeutic procedures for cancer, e.g., planning biopsies or tumour ablations. Manual delineation is extremely labour-intensive, requiring substa…
DiagnosticSemantic SegmentationWeakly supervised Semantic SegmentationWeakly-Supervised Semantic SegmentationTowards localisation of keywords in speech using weak supervision
Developments in weakly supervised and self-supervised models could enable speech technology in low-resource settings where full transcriptions are not available. We consider whether keyword localisation is possible using…
Whose Emotion Matters? Speaking Activity Localisation without Prior Knowledge
The task of emotion recognition in conversations (ERC) benefits from the availability of multiple modalities, as provided, for example, in the video-based Multimodal EmotionLines Dataset (MELD). However, only a few resea…
Active Speaker DetectionAutomatic Speech RecognitionAutomatic Speech Recognition (ASR)Emotion Recognition+2