paper-with-me

홈 › Papers

WhisBERT: Multimodal Text-Audio Language Modeling on 100M Words

2023-12-05 · Lukas Wolf, Greta Tuckute, Klemen Kotar, Eghbal Hosseini, Tamar Regev, Ethan Wilcox, Alex Warstadt

Training on multiple modalities of input can augment the capabilities of a language model. Here, we ask whether such a training regime can improve the quality and efficiency of these systems as well. We focus on text--audio and introduce Whisbert, which is inspired by the text--image approach of FLAVA (Singh et al., 2022). In accordance with Babylm guidelines (Warstadt et al., 2023), we pretrain Whisbert on a dataset comprising only 100 million words plus their corresponding speech from the word-aligned version of the People's Speech dataset (Galvez et al., 2021). To assess the impact of multimodality, we compare versions of the model that are trained on text only and on both audio and text simultaneously. We find that while Whisbert is able to perform well on multimodal masked modeling and surpasses the Babylm baselines in most benchmark tasks, it struggles to optimize its complex objective and outperform its text-only Whisbert baseline.

📄 PDF Abstract BibTeX arXiv:2312.02931

Code (1)

lu-wo/whisbert 공식 구현 pytorch

Tasks

Language ModelingLanguage Modelling

Methods 이 논문이 사용한 방법론

FLAVA FLAVA aims at building a single holistic universal model that targets all modalities at once. FLAVA is a language vision alignment model that learns strong representations from…
Focus 설명 없음

Similar Papers 제목 키워드 기반

UALM: Unified Audio Language Model for Understanding, Generation and Reasoning

2025-10-13 · Jinchuan Tian, Sang-gil Lee, Zhifeng Kong, Sreyan Ghosh 외 arxiv

Recent advances in the audio language modeling (ALM) domain tackle audio understanding and text-to-audio generation as separate tasks. Very few studies attempt to unify these tasks -- an essential step toward advanced mu…

Multimodal ReasoningAudio Generation

CoAVT: A Cognition-Inspired Unified Audio-Visual-Text Pre-Training Model for Multimodal Processing

2024-01-22 · Xianghu Yue, Xiaohai Tian, Lu Lu, Malu Zhang 외

There has been a long-standing quest for a unified audio-visual-text model to enable various multimodal understanding tasks, which mimics the listening, seeing and reading process of human beings. Humans tends to represe…

AudioCapsAudio-Visual SynchronizationLanguage ModelingLanguage Modelling+3

Multilingual Extraction and Recognition of Implicit Discourse Relations in Speech and Text

2026-02-04 · Ahmed Ruby, Christian Hardmeier, Sara Stymne arxiv

Implicit discourse relation classification is a challenging task, as it requires inferring meaning from context. While contextual cues can be distributed across modalities and vary across languages, they are not always c…

Relation ClassificationCross-Lingual Transfer

Make-An-Audio: Text-To-Audio Generation with Prompt-Enhanced Diffusion Models

2023-01-30 · Rongjie Huang, Jiawei Huang, Dongchao Yang, Yi Ren 외

Large-scale multimodal generative modeling has created milestones in text-to-image and text-to-video generation. Its application to audio still lags behind for two main reasons: the lack of large-scale datasets with high…

Audio GenerationText-to-Video GenerationVideo Generation

Multilingual-To-Multimodal (M2M): Unlocking New Languages with Monolingual Text

2026-01-15 · Piyush Singh Pasi arxiv

Multimodal models excel in English, supported by abundant image-text and audio-text data, but performance drops sharply for other languages due to limited multilingual multimodal resources. Existing solutions rely on mac…

Text-to-Image GenerationMachine TranslationImage RetrievalText Retrieval