paper-with-me

홈 › Papers

ViL-Sum: Enhancing Vision and Language Representations via Multi-task Learning for Multi-modal Summarization

2022-01-16 · ACL ARR January 2022 1 · Anonymous

With the advance of multimedia on the Internet, multi-modal summarization has drawn much attention. Most current methods follow a pipeline strategy, where an off-the-shelf object detector is used to extract visual features which are then fused with language representations for decoder to generate. However, these methods suffer two issues 1) separate vision and language representations fail to capture the interrelations within the two modalities; 2) from the local view, the semantic alignments between images and paragraphs are missing. In order to address these problems, in this paper, we propose a novel Vision-Language Summarization (ViL-Sum) model with a multi-task learning framework. Specifically, we train our model with two auxiliary tasks in a multi-task manner, that are images selection and images reordering. In this way, the interrelations within image and text are well captured. Besides, to further enhance the vision-language representation, we employ a unified transformer-based encoder-decoder structure. The encoder simultaneously takes image and text as input and jointly learns the representations of both. Then the representations are used by the decoder to generate the summary. Experimental results show that ViL-Sum significantly outperforms current state-of-the-art methods. In further analysis, we find that the enhanced representations via multi-task training and joint modeling learn reasonable relations between image and text.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

DecoderMulti-Task Learning

Similar Papers 제목 키워드 기반

VL-SAE: Interpreting and Enhancing Vision-Language Alignment with a Unified Concept Set

2025-10-24 · Shufan Shen, Junshu Sun, Qingming Huang, Shuhui Wang arxiv

The alignment of vision-language representations endows current Vision-Language Models (VLMs) with strong multi-modal reasoning capabilities. However, the interpretability of the alignment component remains uninvestigate…

Zero-Shot Image ClassificationSemantic Similarity

X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs

2024-07-18 · Sirnam Swetha, Jinyu Yang, Tal Neiman, Mamshad Nayeem Rizve 외

Recent advancements in Multimodal Large Language Models (MLLMs) have revolutionized the field of vision-language understanding by integrating visual perception capabilities into Large Language Models (LLMs). The prevaili…

Contrastive LearningRepresentation LearningVisual Reasoning

VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization

2026-02-10 · Yikun Liu, Yuan Liu, Shangzhe Di, Haicheng Wang 외 arxiv

Multimodal Large Language Models (MLLMs) have recently achieved remarkable success in visual-language understanding, demonstrating superior high-level semantic alignment within their vision encoders. An important questio…

Semantic SegmentationDepth Estimation

Enhancing Multimodal Large Language Models with Multi-instance Visual Prompt Generator for Visual Representation Enrichment

2024-06-05 · Wenliang Zhong, Wenyi Wu, Qi Li, Rob Barton 외

Multimodal Large Language Models (MLLMs) have achieved SOTA performance in various visual language tasks by fusing the visual representations with LLMs leveraging some visual adapters. In this paper, we first establish t…

Enhancing Generalization in Vision-Language-Action Models by Preserving Pretrained Representations

2025-09-14 · Shresth Grover, Akshay Gopalkrishnan, Bo Ai, Henrik I. Christensen 외 arxiv

Vision-language-action (VLA) models finetuned from vision-language models (VLMs) hold the promise of leveraging rich pretrained representations to build generalist robots across diverse tasks and environments. However, d…

Robot ManipulationSpatial Reasoning