paper-with-me

홈 › Papers

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation

2025-07-10 · Yupu Liang, Yaping Zhang, Zhiyang Zhang, Yang Zhao, Lu Xiang, Chengqing Zong, Yu Zhou arxiv

Document Image Machine Translation (DIMT) aims to translate text within document images, facing generalization challenges due to limited training data and the complex interplay between visual and textual information. To address these challenges, we introduce M4Doc, a novel single-to-mix modality alignment framework leveraging Multimodal Large Language Models (MLLMs). M4Doc aligns an image-only encoder with the multimodal representations of an MLLM, pre-trained on large-scale document image datasets. This alignment enables a lightweight DIMT model to learn crucial visual-textual correlations during training. During inference, M4Doc bypasses the MLLM, maintaining computational efficiency while benefiting from its multimodal knowledge. Comprehensive experiments demonstrate substantial improvements in translation quality, especially in cross-domain generalization and challenging document image scenarios.

📄 PDF Abstract BibTeX arXiv:2507.07572

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyDomain GeneralizationMachine Translation

Similar Papers 제목 키워드 기반

Efficient Remote Sensing with Harmonized Transfer Learning and Modality Alignment

2024-04-28 · Tengjun Huang

With the rise of Visual and Language Pretraining (VLP), an increasing number of downstream tasks are adopting the paradigm of pretraining followed by fine-tuning. Although this paradigm has demonstrated potential in vari…

Cross-Modal RetrievalImage RetrievalImage-to-Text Retrievalparameter-efficient fine-tuning+3

VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

2023-04-17 · Jing Liu, Sihan Chen, Xingjian He, Longteng Guo 외

In this paper, we propose a Vision-Audio-Language Omni-peRception pretraining model (VALOR) for multi-modal understanding and generation. Different from widely-studied vision-language pretraining models, VALOR jointly mo…

Audio captioningAudio-Video Question Answering (AVQA)Audio-Visual CaptioningAudio-visual Question Answering+16

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model

2025-06-16 · Shaolei Zhang, Shoutao Guo, Qingkai Fang, Yan Zhou 외

The emergence of GPT-4o-like large multimodal models (LMMs) has raised the exploration of integrating text, vision, and speech modalities to support more flexible multimodal interaction. Existing LMMs typically concatena…

Large Language Modelmultimodal interaction

MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language Transformer

2024-03-05 · CVPR 2024 1 · JianJian Cao, Peng Ye, Shengze Li, Chong Yu 외

Vision-Language Transformers (VLTs) have shown great success recently, but are meanwhile accompanied by heavy computation costs, where a major reason can be attributed to the large number of visual and language tokens. E…

How to Bridge the Gap between Modalities: Survey on Multimodal Large Language Model

2023-11-10 · Shezheng Song, Xiaopeng Li, Shasha Li, Shan Zhao 외

We explore Multimodal Large Language Models (MLLMs), which integrate LLMs like GPT-4 to handle multimodal data, including text, images, audio, and more. MLLMs demonstrate capabilities such as generating image captions an…

Image CaptioningLanguage ModelingLanguage ModellingLarge Language Model+1