paper-with-me

Papers

Make Imagination Clearer! Stable Diffusion-based Visual Imagination for Multimodal Machine Translation

2024-12-17 · Andong Chen, Yuchen Song, Kehai Chen, Muyun Yang, Tiejun Zhao, Min Zhang

Visual information has been introduced for enhancing machine translation (MT), and its effectiveness heavily relies on the availability of large amounts of bilingual parallel sentence pairs with manual image annotations. In this paper, we introduce a stable diffusion-based imagination network into a multimodal large language model (MLLM) to explicitly generate an image for each source sentence, thereby advancing the multimodel MT. Particularly, we build heuristic human feedback with reinforcement learning to ensure the consistency of the generated image with the source sentence without the supervision of image annotation, which breaks the bottleneck of using visual information in MT. Furthermore, the proposed method enables imaginative visual information to be integrated into large-scale text-only MT in addition to multimodal MT. Experimental results show that our model significantly outperforms existing multimodal MT and text-only MT, especially achieving an average improvement of more than 14 BLEU points on Multi30K multimodal MT benchmarks.

📄 PDF Abstract BibTeX arXiv:2412.12627

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language ModelMachine TranslationMultimodal Large Language ModelMultimodal Machine TranslationSentence

Similar Papers 제목 키워드 기반

ImaginE: An Imagination-Based Automatic Evaluation Metric for Natural Language Generation

2021-06-10 · Wanrong Zhu, Xin Eric Wang, An Yan, Miguel Eckstein 외

Automatic evaluations for natural language generation (NLG) conventionally rely on token-level or embedding-level comparisons with text references. This differs from human language processing, for which visual imaginatio…

nlg evaluationText Generation

Dreaming the Unseen: World Model-regularized Diffusion Policy for Out-of-Distribution Robustness

2026-03-22 · Ziou Hu, Xiangtong Yao, Yuan Meng, Zhenshan Bing 외 arxiv

Diffusion policies excel at visuomotor control but often fail catastrophically under severe out-of-distribution (OOD) disturbances, such as unexpected object displacements or visual corruptions. To address this vulnerabi…

Learning to Imagine: Visually-Augmented Natural Language Generation

2023-05-26 · Tianyi Tang, Yushuo Chen, Yifan Du, Junyi Li 외

People often imagine relevant scenes to aid in the writing process. In this work, we aim to utilize visual information for composition in the same manner as humans. We propose a method, LIVE, that makes pre-trained langu…

SentenceText Generation

Stale Diffusion: Hyper-realistic 5D Movie Generation Using Old-school Methods

2024-04-01 · Joao F. Henriques, Dylan Campbell, Tengda Han

Two years ago, Stable Diffusion achieved super-human performance at generating images with super-human numbers of fingers. Following the steady decline of its technical novelty, we propose Stale Diffusion, a method that …

Do Visual Imaginations Improve Vision-and-Language Navigation Agents?

2025-03-20 · CVPR 2025 1 · Akhil Perincherry, Jacob Krantz, Stefan Lee

Vision-and-Language Navigation (VLN) agents are tasked with navigating an unseen environment using natural language instructions. In this work, we study if visual representations of sub-goals implied by the instructions …

Vision and Language Navigation