paper-with-me

Papers

Sequence-to-Sequence Multi-Modal Speech In-Painting

2024-06-03 · Mahsa Kadkhodaei Elyaderani, Shahram Shirani

Speech in-painting is the task of regenerating missing audio contents using reliable context information. Despite various recent studies in multi-modal perception of audio in-painting, there is still a need for an effective infusion of visual and auditory information in speech in-painting. In this paper, we introduce a novel sequence-to-sequence model that leverages the visual information to in-paint audio signals via an encoder-decoder architecture. The encoder plays the role of a lip-reader for facial recordings and the decoder takes both encoder outputs as well as the distorted audio spectrograms to restore the original speech. Our model outperforms an audio-only speech in-painting model and has comparable results with a recent multi-modal speech in-painter in terms of speech quality and intelligibility metrics for distortions of 300 ms to 1500 ms duration, which proves the effectiveness of the introduced multi-modality in speech in-painting.

📄 PDF Abstract BibTeX arXiv:2406.01321

Code (0)

등록된 구현이 없습니다.

Tasks

Decoder

Similar Papers 제목 키워드 기반

Robust Multi-Modal Speech In-Painting: A Sequence-to-Sequence Approach

2024-06-02 · Mahsa Kadkhodaei Elyaderani, Shahram Shirani

The process of reconstructing missing parts of speech audio from context is called speech in-painting. Human perception of speech is inherently multi-modal, involving both audio and visual (AV) cues. In this paper, we in…

Lip ReadingMulti-Task Learning

STEMM: Self-learning with Speech-text Manifold Mixup for Speech Translation

2022-03-20 · ACL 2022 5 · Qingkai Fang, Rong Ye, Lei LI, Yang Feng 외

How to learn a better speech representation for end-to-end speech-to-text translation (ST) with limited labeled data? Existing techniques often attempt to transfer powerful machine translation (MT) capabilities to ST, bu…

Machine TranslationSpeech-to-TextSpeech-to-Text Translation

VQ-CTAP: Cross-Modal Fine-Grained Sequence Representation Learning for Speech Processing

2024-08-11 · Chunyu Qiang, Wang Geng, Yi Zhao, Ruibo Fu 외

Deep learning has brought significant improvements to the field of cross-modal representation learning. For tasks such as text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), a cross-modal…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Representation Learningspeech-recognition+4

Temporal Multimodal Learning in Audiovisual Speech Recognition

2016-06-01 · CVPR 2016 6 · Di Hu, Xuelong. Li, Xiaoqiang Lu

In view of the advantages of deep networks in producing useful representation, the generated features of different modality data (such as image, audio) can be jointly learned using Multimodal Restricted Boltzmann Machine…

Multimodal Deep Learningspeech-recognitionSpeech Recognition

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

2025-06-16 · Shaolei Zhang, Shoutao Guo, Qingkai Fang, Yan Zhou 외

The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal interaction. Existing LMMs typically concatena…

Large Language Modelmultimodal interaction