paper-with-me

Papers

Fused Acoustic and Text Encoding for Multimodal Bilingual Pretraining and Speech Translation

2021-02-10 · Renjie Zheng, Junkun Chen, Mingbo Ma, Liang Huang

Recently, representation learning for text and speech has successfully improved many language related tasks. However, all existing methods suffer from two limitations: (a) they only learn from one input modality, while a unified representation for both speech and text is needed by tasks such as end-to-end speech translation, and as a result,(b) they can not exploit various large-scale text and speech data and their performance is limited by the scarcity of parallel speech translation data.To address these problems, we propose a Fused Acoustic and Text Masked Language Model (FAT-MLM) which jointly learns a unified representation for both acoustic and text input from various types of corpora including parallel data for speech recognition and machine translation, and even pure speech and text data. Within this cross-modal representation learning framework, we further present an end-to-end model for Fused Acoustic and Text Speech Translation (FAT-ST). Experiments on three translation directions show that by fine-tuning from FAT-MLM, our proposed speech translation models substantially improve translation quality by up to +5.9 BLEU.

📄 PDF Abstract BibTeX arXiv:2102.05766

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingMachine TranslationRepresentation Learningspeech-recognitionSpeech RecognitionTranslation

Similar Papers 제목 키워드 기반

Revisiting Cross Modal Retrieval

2018-07-19 · Shah Nawaz, Muhammad Kamran Janjua, Alessandro Calefati, Ignazio Gallo

This paper proposes a cross-modal retrieval system that leverages on image and text encoding. Most multimodal architectures employ separate networks for each modality to capture the semantic relationship between them. Ho…

Cross-Modal RetrievalRetrieval

Beyond the Baseband: Adaptive Multi-Band Encoding for Full-Spectrum Bioacoustics Classification

2026-04-30 · Eklavya Sarkar, Marius Miron, David Robinson, Gagan Narula 외 arxiv

Animals hear and vocalize across frequency ranges that differ substantially from humans, often extending into the ultrasonic domain. Yet most computational bioacoustics systems rely on audio models pre-trained at 16 kHz,…

Taiyi-Diffusion-XL: Advancing Bilingual Text-to-Image Generation with Large Vision-Language Model Support

2024-01-26 · XiaoJun Wu, Dixiang Zhang, Ruyi Gan, Junyu Lu 외

Recent advancements in text-to-image models have significantly enhanced image generation capabilities, yet a notable gap of open-source models persists in bilingual or Chinese language support. To address this need, we p…

Image GenerationLanguage ModelingLanguage ModellingText to Image Generation+1

A Multimodal Framework for Dementia Detection via Linguistic and Acoustic Representation Learning

2026-05-25 · Loukas Ilias, Dimitris Askounis arxiv

Alzheimer's disease (AD) is a progressive neurodegenerative disorder and the leading cause of dementia, affecting memory, reasoning, communication, and daily functioning. Early diagnosis is particularly important, as tim…

Multimodal Deep LearningRepresentation Learning

Speech-Guided Multimodal Learning for Vocal Tract Segmentation in Real-Time MRI

2026-05-18 · Daiqi Liu, Lukas Mulzer, Md Hasan, Nyvenn de Castro 외 arxiv

Segmenting vocal tract articulators in real-time MRI (rtMRI) is a challenging dynamic image segmentation problem characterized by low contrast, rapid motion, and limited spatial resolution. However, while rtMRI acquisiti…

Image Segmentation