paper-with-me

Papers

VATLM: Visual-Audio-Text Pre-Training with Unified Masked Prediction for Speech Representation Learning

2022-11-21 · Qiushi Zhu, Long Zhou, Ziqiang Zhang, Shujie Liu, Binxing Jiao, Jie Zhang, LiRong Dai, Daxin Jiang, Jinyu Li, Furu Wei

Although speech is a simple and effective way for humans to communicate with the outside world, a more realistic speech interaction contains multimodal information, e.g., vision, text. How to design a unified framework to integrate different modal information and leverage different resources (e.g., visual-audio pairs, audio-text pairs, unlabeled speech, and unlabeled text) to facilitate speech representation learning was not well explored. In this paper, we propose a unified cross-modal representation learning framework VATLM (Visual-Audio-Text Language Model). The proposed VATLM employs a unified backbone network to model the modality-independent information and utilizes three simple modality-dependent modules to preprocess visual, speech, and text inputs. In order to integrate these three modalities into one shared semantic space, VATLM is optimized with a masked prediction task of unified tokens, given by our proposed unified tokenizer. We evaluate the pre-trained VATLM on audio-visual related downstream tasks, including audio-visual speech recognition (AVSR), visual speech recognition (VSR) tasks. Results show that the proposed VATLM outperforms previous the state-of-the-art models, such as audio-visual pre-trained AV-HuBERT model, and analysis also demonstrates that VATLM is capable of aligning different modalities into the same space. To facilitate future research, we release the code and pre-trained models at https://aka.ms/vatlm.

📄 PDF Abstract BibTeX arXiv:2211.11275

Code (0)

등록된 구현이 없습니다.

Tasks

Audio-Visual Speech RecognitionLanguage ModellingRepresentation Learningspeech-recognitionSpeech RecognitionSpeech Representation LearningVisual Speech Recognition

Similar Papers 제목 키워드 기반

From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

2024-09-27 · Kun Su, Xiulong Liu, Eli Shlizerman

Video encompasses both visual and auditory data, creating a perceptually rich experience where these two modalities complement each other. As such, videos are a valuable type of media for the investigation of the interpl…

Audio ClassificationAudio GenerationRepresentation Learning

CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing

2024-01-22 · Xianghu Yue, Xiaohai Tian, Lu Lu, Malu Zhang 외

There has been a long-standing quest for a unified audio-visual-text model to enable various multimodal understanding tasks, which mimics the listening, seeing and reading process of human beings. Humans tends to represe…

AudioCapsAudio-Visual SynchronizationLanguage ModelingLanguage Modelling+3

Unified Video-Language Pre-training with Synchronized Audio

2024-05-12 · Shentong Mo, Haofan Wang, Huaxia Li, Xu Tang

Video-language pre-training is a typical and challenging problem that aims at learning visual and textual representations from large-scale data in a self-supervised way. Existing pre-training approaches either captured t…

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling

2026-04-16 · Jianxuan Yang, Xinyue Guo, Zhi Cheng, Kai Wang 외 arxiv

Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-grained controllability remains challenging. Existing methods suffer from weak tex…

Audio Generation

Sounding Video Generator: A Unified Framework for Text-guided Sounding Video Generation

2023-03-29 · Jiawei Liu, Weining Wang, Sihan Chen, Xinxin Zhu 외

As a combination of visual and audio signals, video is inherently multi-modal. However, existing video generation methods are primarily intended for the synthesis of visual frames, whereas audio signals in realistic vide…

Audio GenerationContrastive LearningDecoderVideo Generation