paper-with-me

Papers

VLIS: Unimodal Language Models Guide Multimodal Language Generation

2023-10-15 · Jiwan Chung, Youngjae Yu

Multimodal language generation, which leverages the synergy of language and vision, is a rapidly expanding field. However, existing vision-language models face challenges in tasks that require complex linguistic understanding. To address this issue, we introduce Visual-Language models as Importance Sampling weights (VLIS), a novel framework that combines the visual conditioning capability of vision-language models with the language understanding of unimodal text-only language models without further training. It extracts pointwise mutual information of each image and text from a visual-language model and uses the value as an importance sampling weight to adjust the token likelihood from a text-only model. VLIS improves vision-language models on diverse tasks, including commonsense understanding (WHOOPS, OK-VQA, and ScienceQA) and complex text generation (Concadia, Image Paragraph Captioning, and ROCStories). Our results suggest that VLIS represents a promising new direction for multimodal language generation.

📄 PDF Abstract BibTeX arXiv:2310.09767

Code (1)

jiwanchung/vlis 공식 구현 pytorch

Tasks

Caption GenerationExplanation GenerationImage Paragraph CaptioningLanguage ModelingLanguage ModellingText GenerationVisual Question Answering (VQA)Zero-Shot Image Paragraph Captioning

Similar Papers 제목 키워드 기반

MM-SHAP: A Performance-agnostic Metric for Measuring Multimodal Contributions in Vision and Language Models & Tasks

2022-12-15 · Letitia Parcalabescu, Anette Frank

Vision and language models (VL) are known to exploit unrobust indicators in individual modalities (e.g., introduced by distributional biases) instead of focusing on relevant information in each modality. That a unimodal …

Quantifying and Mitigating Unimodal Biases in Multimodal Large Language Models: A Causal Perspective

2024-03-27 · Meiqi Chen, Yixin Cao, Yan Zhang, Chaochao Lu

Recent advancements in Large Language Models (LLMs) have facilitated the development of Multimodal LLMs (MLLMs). Despite their impressive capabilities, MLLMs often suffer from over-reliance on unimodal biases (e.g., lang…

Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Instruction-Tuned Video-Audio Models Elucidate Functional Specialization in the Brain

2025-06-09 · Subba Reddy Oota, Khushbu Pahwa, Prachi Jindal, Satya Sai Srinath Namburi 외

Recent voxel-wise multimodal brain encoding studies have shown that multimodal large language models (MLLMs) exhibit a higher degree of brain alignment compared to unimodal models in both unimodal and multimodal stimulus…

Disentanglement

Meta-Learn Unimodal Signals with Weak Supervision for Multimodal Sentiment Analysis

2024-08-28 · Sijie Mai, Yu Zhao, Ying Zeng, Jianhua Yao 외

Multimodal sentiment analysis aims to effectively integrate information from various sources to infer sentiment, where in many cases there are no annotations for unimodal labels. Therefore, most works rely on multimodal …

DenoisingMultimodal Sentiment AnalysisSentiment Analysis

MemoSen: A Multimodal Dataset for Sentiment Analysis of Memes

2022-06-01 · LREC 2022 6 · Eftekhar Hossain, Omar Sharif, Mohammed Moshiul Hoque

Posting and sharing memes have become a powerful expedient of expressing opinions on social media in recent days. Analysis of sentiment from memes has gained much attention to researchers due to its substantial implicati…

Multimodal Sentiment AnalysisSentiment AnalysisSentiment Classification