paper-with-me

홈 › Papers

Multimodal Machine Translation through Visuals and Speech

2019-11-28 · Umut Sulubacak, Ozan Caglayan, Stig-Arne Grönroos, Aku Rouhe, Desmond Elliott, Lucia Specia, Jörg Tiedemann

Multimodal machine translation involves drawing information from more than one modality, based on the assumption that the additional modalities will contain useful alternative views of the input data. The most prominent tasks in this area are spoken language translation, image-guided translation, and video-guided translation, which exploit audio and visual modalities, respectively. These tasks are distinguished from their monolingual counterparts of speech recognition, image captioning, and video captioning by the requirement of models to generate outputs in a different language. This survey reviews the major data resources for these tasks, the evaluation campaigns concentrated around them, the state of the art in end-to-end and pipeline approaches, and also the challenges in performance evaluation. The paper concludes with a discussion of directions for future research in these areas: the need for more expansive and challenging datasets, for targeted evaluations of model performance, and for multimodality in both the input and output space.

📄 PDF Abstract BibTeX arXiv:1911.12798

Code (0)

등록된 구현이 없습니다.

Tasks

Image CaptioningMachine TranslationMultimodal Machine Translationspeech-recognitionSpeech RecognitionTranslationVideo Captioning

Similar Papers 제목 키워드 기반

The Art of Storytelling: Multi-Agent Generative AI for Dynamic Multimodal Narratives

2024-09-17 · Samee Arif, Taimoor Arif, Muhammad Saad Haroon, Aamina Jamal Khan 외

This paper introduces the concept of an education tool that utilizes Generative Artificial Intelligence (GenAI) to enhance storytelling for children. The system combines GenAI-driven narrative co-creation, text-to-speech…

text-to-speechText to SpeechText-to-Video GenerationVideo Generation

Scalable Multilingual Multimodal Machine Translation with Speech-Text Fusion

2026-02-25 · Yexing Du, Youcheng Pan, Zekun Wang, Zheng Chu 외 arxiv

Multimodal Large Language Models (MLLMs) have achieved notable success in enhancing translation performance by integrating multimodal information. However, existing research primarily focuses on image-guided methods, who…

Multimodal Machine Translation

SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

2023-08-22 · Seamless Communication, Loïc Barrault, Yu-An Chung, Mariano Cora Meglioli 외

What does it take to create the Babel Fish, a tool that can help individuals translate speech between any two languages? While recent breakthroughs in text-based models have pushed machine translation coverage beyond 200…

Automatic Speech RecognitionMachine TranslationSpeech-to-Speech TranslationSpeech-to-Text+5

EMMeTT: Efficient Multimodal Machine Translation Training

2024-09-20 · Piotr Żelasko, Zhehuai Chen, Mengru Wang, Daniel Galvez 외

A rising interest in the modality extension of foundation language models warrants discussion on the most effective, and efficient, multimodal training approach. This work focuses on neural machine translation (NMT) and …

automatic-speech-translationDecoderMachine TranslationMultimodal Machine Translation+2

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

2026-04-13 · Bingzheng Qu, Kehai Chen, Xuefeng Bai, Min Zhang arxiv

Recent progress in multimodal large language models (MLLMs) is reshaping video translation from a cascaded pipeline of automatic speech recognition, machine translation, text-to-speech, and lip synchronization into a uni…

Multimodal ReasoningMachine TranslationSpeech Recognition