paper-with-me

홈 › Papers

The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning

2024-11-18 · Longju Bai, Angana Borah, Oana Ignat, Rada Mihalcea

Large Multimodal Models (LMMs) exhibit impressive performance across various multimodal tasks. However, their effectiveness in cross-cultural contexts remains limited due to the predominantly Western-centric nature of most data and models. Conversely, multi-agent models have shown significant capability in solving complex tasks. Our study evaluates the collective performance of LMMs in a multi-agent interaction setting for the novel task of cultural image captioning. Our contributions are as follows: (1) We introduce MosAIC, a Multi-Agent framework to enhance cross-cultural Image Captioning using LMMs with distinct cultural personas; (2) We provide a dataset of culturally enriched image captions in English for images from China, India, and Romania across three datasets: GeoDE, GD-VCR, CVQA; (3) We propose a culture-adaptable metric for evaluating cultural information within image captions; and (4) We show that the multi-agent interaction outperforms single-agent models across different metrics, and offer valuable insights for future research. Our dataset and models can be accessed at https://github.com/MichiganNLP/MosAIC.

📄 PDF Abstract BibTeX arXiv:2411.11758

Code (1)

michigannlp/mosaic 공식 구현 pytorch

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

AgroTools: A Benchmark for Tool-Augmented Multimodal Agents in Agriculture

2026-05-21 · Zi Ye, Yibin Wen, Xiaoya Fan, Xinyu Zhang 외 arxiv

Agricultural decision-making increasingly requires multimodal systems that can transform visual observations into reliable, executable actions. However, existing agricultural multimodal benchmarks mainly evaluate final-a…

Multi-Agent Multimodal Models for Multicultural Text to Image Generation

2025-02-21 · Parth Bhalerao, Mounika Yalamarty, Brian Trinh, Oana Ignat

Large Language Models (LLMs) demonstrate impressive performance across various multimodal tasks. However, their effectiveness in cross-cultural contexts remains limited due to the predominantly Western-centric nature of …

Image GenerationText to Image GenerationText-to-Image Generation

CM2: Multimodal Cultural Reasoning via an Integrated Multi-Agent Framework

2026-08-31 · Qi Li, Zhaojie Kang, Yingjie He, Zheng Lin 외 arxiv

Multimodal Large Language Models (MLLMs) have shown remarkable success in STEM domains, where progress is often driven by vertical, step-by-step deduction under relatively stable symbol systems. Their horizontal, interdi…

VietMEAgent: Culturally-Aware Few-Shot Multimodal Explanation for Vietnamese Visual Question Answering

2025-11-12 · Hai-Dang Nguyen, Minh-Anh Dang, Minh-Tan Le, Minh-Tuan Le arxiv

Contemporary Visual Question Answering (VQA) systems remain constrained when confronted with culturally specific content, largely because cultural knowledge is under-represented in training corpora and the reasoning proc…

Visual Question AnsweringObject Detection

LLandMark: A Multi-Agent Framework for Landmark-Aware Multimodal Interactive Video Retrieval

2026-03-03 · Minh-Chi Phung, Thien-Bao Le, Cam-Tu Tran-Thi, Thu-Dieu Nguyen-Thi 외 arxiv

The increasing diversity and scale of video data demand retrieval systems capable of multimodal understanding, adaptive reasoning, and domain-specific knowledge integration. This paper presents LLandMark, a modular multi…

Video Retrieval