paper-with-me

홈 › Papers

A Multimodal Recaptioning Framework to Account for Perceptual Diversity in Multilingual Vision-Language Modeling

2025-04-19 · Kyle Buettner, Jacob Emmerson, Adriana Kovashka

There are many ways to describe, name, and group objects when captioning an image. Differences are evident when speakers come from diverse cultures due to the unique experiences that shape perception. Machine translation of captions has pushed multilingual capabilities in vision-language models (VLMs), but data comes mainly from English speakers, indicating a perceptual bias and lack of model flexibility. In this work, we address this challenge and outline a data-efficient framework to instill multilingual VLMs with greater understanding of perceptual diversity. We specifically propose an LLM-based, multimodal recaptioning strategy that alters the object descriptions of English captions before translation. The greatest benefits are demonstrated in a targeted multimodal mechanism guided by native speaker data. By adding produced rewrites as augmentations in training, we improve on German and Japanese text-image retrieval cases studies (up to +3.5 mean recall overall, +4.7 on non-native error cases). We further propose a mechanism to analyze the specific object description differences across datasets, and we offer insights into cross-dataset and cross-language generalization.

📄 PDF Abstract BibTeX arXiv:2504.14359

Code (0)

등록된 구현이 없습니다.

Tasks

DiversityImage RetrievalLanguage ModelingLanguage ModellingMachine TranslationTranslation

Similar Papers 제목 키워드 기반

Perceptual Rationality: An Evolutionary Game Theory of Perceptually Rational Decision-Making

2025-06-21 · Mohammad Salahshour

Understanding how biological organisms make decisions is of fundamental importance in understanding behavior. Such an understanding within evolutionary game theory so far has been sought by appealing to bounded rationali…

Decision MakingDiversity

Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs

2024-01-22 · Ling Yang, Zhaochen Yu, Chenlin Meng, Minkai Xu 외

Diffusion models have exhibit exceptional performance in text-to-image generation and editing. However, existing methods often face challenges when handling complex text prompts that involve multiple objects with multipl…

Diffusion Personalization Tuning FreeImage GenerationLarge Language ModelText to Image Generation+1

When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy

2026-05-12 · Xiaofeng Tan, Jun Liu, Bin-Bin Gao, Yuanting Fan 외 arxiv

RLHF is widely used to align flow-matching text-to-image models with human preferences, but often leads to severe diversity collapse after fine-tuning. In RL, diversity is often assumed to correlate with policy entropy, …

Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis

2026-03-31 · Shuang Chen, Quanxin Shou, Hangting Chen, Yucheng Zhou 외 arxiv

Unified multimodal models provide a natural and promising architecture for understanding diverse and complex real-world knowledge while generating high-quality images. However, they still rely primarily on frozen paramet…

Image Generation

TAMP: Token-Adaptive Layerwise Pruning in Multimodal Large Language Models

2025-04-14 · Jaewoo Lee, Keyang Xuan, Chanakya Ekbote, Sandeep Polisetty 외

Multimodal Large Language Models (MLLMs) have shown remarkable versatility in understanding diverse multimodal data and tasks. However, these capabilities come with an increased model scale. While post-training pruning r…

Diversity