ViDove: A Translation Agent System with Multimodal Context and Memory-Augmented Reasoning
LLM-based translation agents have achieved highly human-like translation results and are capable of handling longer and more complex contexts with greater efficiency. However, they are typically limited to text-only inputs. In this paper, we introduce ViDove, a translation agent system designed for multimodal input. Inspired by the workflow of human translators, ViDove leverages visual and contextual background information to enhance the translation process. Additionally, we integrate a multimodal memory system and long-short term memory modules enriched with domain-specific knowledge, enabling the agent to perform more accurately and adaptively in real-world scenarios. As a result, ViDove achieves significantly higher translation quality in both subtitle generation and general translation tasks, with a 28% improvement in BLEU scores and a 15% improvement in SubER compared to previous state-of-the-art baselines. Moreover, we introduce DoveBench, a new benchmark for long-form automatic video subtitling and translation, featuring 17 hours of high-quality, human-annotated data. Our code is available here: https://github.com/pigeonai-org/ViDove
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
MSCTD: A Multimodal Sentiment Chat Translation Dataset
Multimodal machine translation and textual chat translation have received considerable attention in recent years. Although the conversation in its natural form is usually multimodal, there still lacks work on multimodal …
Machine TranslationMultimodal Machine TranslationSentiment AnalysisTranslationLIUM-CVC Submissions for WMT17 Multimodal Translation Task
This paper describes the monomodal and multimodal Neural Machine Translation systems developed by LIUM and CVC for WMT17 Shared Task on Multimodal Translation. We mainly explored two multimodal architectures where either…
Machine TranslationTranslationCaMMT: Benchmarking Culturally Aware Multimodal Machine Translation
Cultural content poses challenges for machine translation systems due to the differences in conceptualizations between cultures, where language alone may fail to convey sufficient context to capture region-specific meani…
BenchmarkingMachine TranslationMultimodal Machine TranslationTranslationExploiting Multimodal Reinforcement Learning for Simultaneous Machine Translation
This paper addresses the problem of simultaneous machine translation (SiMT) by exploring two main concepts: (a) adaptive policies to learn a good trade-off between high translation quality and low latency; and (b) visual…
Machine Translationreinforcement-learningReinforcement LearningReinforcement Learning (RL)+1Impact of Visual Context on Noisy Multimodal NMT: An Empirical Study for English to Indian Languages
The study investigates the effectiveness of utilizing multimodal information in Neural Machine Translation (NMT). While prior research focused on using multimodal data in low-resource scenarios, this study examines how i…
Machine TranslationNMTTranslation