paper-with-me

Papers

Cued Speech Generation Leveraging a Pre-trained Audiovisual Text-to-Speech Model

2025-01-08 · Sanjana Sankar, Martin Lenglet, Gerard Bailly, Denis Beautemps, Thomas Hueber

This paper presents a novel approach for the automatic generation of Cued Speech (ACSG), a visual communication system used by people with hearing impairment to better elicit the spoken language. We explore transfer learning strategies by leveraging a pre-trained audiovisual autoregressive text-to-speech model (AVTacotron2). This model is reprogrammed to infer Cued Speech (CS) hand and lip movements from text input. Experiments are conducted on two publicly available datasets, including one recorded specifically for this study. Performance is assessed using an automatic CS recognition system. With a decoding accuracy at the phonetic level reaching approximately 77%, the results demonstrate the effectiveness of our approach.

📄 PDF Abstract BibTeX arXiv:2501.04799

Code (0)

등록된 구현이 없습니다.

Tasks

text-to-speechText to SpeechTransfer Learning

Similar Papers 제목 키워드 기반

Enhancing Audiovisual Speech Recognition through Bifocal Preference Optimization

2024-12-26 · Yihan Wu, Yichen Lu, Yifan Peng, Xihua Wang 외

Audiovisual Automatic Speech Recognition (AV-ASR) aims to improve speech recognition accuracy by leveraging visual signals. It is particularly challenging in unconstrained real-world scenarios across various domains due …

Automatic Speech Recognitionspeech-recognitionSpeech Recognition

Robust Audiovisual Speech Recognition Models with Mixture-of-Experts

2024-09-19 · Yihan Wu, Yifan Peng, Yichen Lu, Xuankai Chang 외

Visual signals can enhance audiovisual speech recognition accuracy by providing additional contextual information. Given the complexity of visual signals, an audiovisual speech recognition model requires robust generaliz…

Mixture-of-ExpertsRobust Speech Recognitionspeech-recognitionSpeech Recognition

Investigations on End-to-End Audiovisual Fusion

2018-04-30 · Michael Wand, Ngoc Thang Vu, Juergen Schmidhuber

Audiovisual speech recognition (AVSR) is a method to alleviate the adverse effect of noise in the acoustic signal. Leveraging recent developments in deep neural network-based speech recognition, we present an AVSR neural…

speech-recognitionSpeech Recognition

AV-CrossNet: an Audiovisual Complex Spectral Mapping Network for Speech Separation By Leveraging Narrow- and Cross-Band Modeling

2024-06-17 · Vahid Ahmadi Kalkhorani, Cheng Yu, Anurag Kumar, Ke Tan 외

Adding visual cues to audio-based speech separation can improve separation performance. This paper introduces AV-CrossNet, an \gls{av} system for speech enhancement, target speaker extraction, and multi-talker speaker se…

Speaker SeparationSpeech EnhancementSpeech SeparationTarget Speaker Extraction

A vector quantized masked autoencoder for audiovisual speech emotion recognition

2023-05-05 · Samir Sadok, Simon Leglaive, Renaud Séguier

An important challenge in emotion recognition is to develop methods that can leverage unlabeled training data. In this paper, we propose the VQ-MAE-AV model, a self-supervised multimodal model that leverages masked autoe…

Contrastive LearningEmotion RecognitionRepresentation LearningSelf-Supervised Learning+1