paper-with-me

홈 › Papers

VLURes: Benchmarking VLM Visual and Linguistic Understanding in Low-Resource Languages

2025-10-14 · Jesse Atuhurra, Iqra Ali, Tomoya Iwakura, Hidetaka Kamigaito, Tatsuya Hiraoka arxiv

Vision Language Models (VLMs) are pivotal for advancing perception in intelligent agents. Yet, evaluation of VLMs remains limited to predominantly English-centric benchmarks in which the image-text pairs comprise short texts. To evaluate VLM fine-grained abilities, in four languages under long-text settings, we introduce a novel multilingual benchmark VLURes featuring eight vision-and-language tasks, and a pioneering unrelatedness task, to probe the fine-grained Visual and Linguistic Understanding capabilities of VLMs across English, Japanese, and low-resource languages, Swahili, and Urdu. Our datasets, curated from web resources in the target language, encompass ten diverse image categories and rich textual context, introducing valuable vision-language resources for Swahili and Urdu. By prompting VLMs to generate responses and rationales, evaluated automatically and by native speakers, we uncover performance disparities across languages and tasks critical to intelligent agents, such as object recognition, scene understanding, and relationship understanding. We conducted evaluations of ten VLMs with VLURes. The best performing model, GPT-4o, achieves an overall accuracy of 90.8% and lags human performance by 6.7%, though the gap is larger for open-source models. The gap highlights VLURes' critical role in developing intelligent agents to tackle multi-modal visual reasoning.

📄 PDF Abstract BibTeX arXiv:2510.12845

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingObject RecognitionVisual Reasoning

Similar Papers 제목 키워드 기반

BLEnD-Vis: Benchmarking Multimodal Cultural Understanding in Vision Language Models

2025-10-13 · Bryan Chen Zhengyu Tan, Zheng Weihua, Zhengyuan Liu, Nancy F. Chen 외 arxiv

As vision-language models (VLMs) are deployed globally, their ability to understand culturally situated knowledge becomes essential. Yet, existing evaluations largely assess static recall or isolated visual grounding, le…

Visual Grounding

Visualizing Linguistic Diversity of Text Datasets Synthesized by Large Language Models

2023-05-19 · Emily Reif, Minsuk Kahng, Savvas Petridis

Large language models (LLMs) can be used to generate smaller, more refined datasets via few-shot prompting for benchmarking, fine-tuning or other use cases. However, understanding and evaluating these datasets is difficu…

BenchmarkingDiversity

AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering

2026-03-10 · Nguyen Anh Tuong, Phan Ba Duc, Nguyen Trung Quoc, Tran Dac Thinh 외 arxiv

Visual Question Answering (VQA) is a fundamental multimodal task that requires models to jointly understand visual and textual information. Early VQA systems relied heavily on language biases, motivating subsequent work …

Visual Question AnsweringRepresentation LearningMachine TranslationImage Captioning

Benchmarking Vision Language Models for Cultural Understanding

2024-07-15 · Shravan Nayak, Kanishk Jain, Rabiul Awal, Siva Reddy 외

Foundation models and vision-language pre-training have notably advanced Vision Language Models (VLMs), enabling multimodal processing of visual and linguistic data. However, their performance has been typically assessed…

BenchmarkingQuestion AnsweringScene UnderstandingVisual Question Answering

LLM Probe: Evaluating LLMs for Low-Resource Languages

2026-03-31 · Hailay Kidu Teklehaymanot, Gebrearegawi Gebremariam, Wolfgang Nejdl arxiv

Despite rapid advances in large language models (LLMs), their linguistic abilities in low-resource and morphologically rich languages are still not well understood due to limited annotated resources and the absence of st…

Speech Recognition