paper-with-me

Papers

Not Truly Multilingual: Script Consistency as a Missing Dimension in VLM Evaluation

2026-06-15 · Prabhjot Singh, Bhushan Pawar, Madhu Reddiboina, Rajvee Sheth arxiv

Current multilingual evaluations for Vision-Language Models (VLMs) assume a one-to-one mapping between language and orthography, overlooking billions of users of multi-script languages. We introduce PuMVR (Punjabi Multimodal Visual Reasoning), a benchmark of 1,000 strictly parallel image-text instances across Punjabi's three active scripts: Gurmukhi, Shahmukhi, and Roman. Evaluating 10 state-of-the-art VLMs, we expose a substantial and systematic Script Gap. Models frequently solve visual tasks in one script while failing identical tasks in another, with accuracy deltas reaching 16%. Crucially, visual input boosts absolute performance uniformly yet does not close the orthographic gap. Furthermore, cross-script in-context transfer is highly brittle, exposing script-locked knowledge representation. Supported by McNemar tests across all script pairs, our findings demonstrate that current "multilingual" VLMs are not truly multi-script. We propose the Script Consistency Rate (SCR), which falls as low as 24.8% on our benchmark, as a mandatory metric for script-agnostic evaluation to ensure equitable AI access. Data and code are available at: https://github.com/prabhjotschugh/Not-Truly-Multilingual-PuMVR.

📄 PDF Abstract BibTeX arXiv:2606.17188

Code (0)

등록된 구현이 없습니다.

Tasks

Visual Reasoning

Similar Papers 제목 키워드 기반

Found in Translation: Measuring Multilingual LLM Consistency as Simple as Translate then Evaluate

2025-05-28 · Ashim Gupta, Maitrey Mehta, Zhichao Xu, Vivek Srikumar

Large language models (LLMs) provide detailed and impressive responses to queries in English. However, are they really consistent at responding to the same query in other languages? The popular way of evaluating for mult…

Benchmarking

A Weakly-Supervised Streaming Multilingual Speech Model with Truly Zero-Shot Capability

2022-11-04 · Jian Xue, Peidong Wang, Jinyu Li, Eric Sun

In this paper, we introduce our work of building a Streaming Multilingual Speech Model (SM2), which can transcribe or translate multiple spoken languages into texts of the target language. The backbone of SM2 is Transfor…

Machine Translationspeech-recognitionSpeech RecognitionTranslation

Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs

2025-05-21 · Hao Wang, Pinzhi Huang, Jihan Yang, Saining Xie 외

The rapid evolution of multimodal large language models (MLLMs) has significantly enhanced their real-world applications. However, achieving consistent performance across languages, especially when integrating cultural k…

BenchmarkingQuestion AnsweringVisual Question Answering

Hybrid Approximate Nearest Neighbor Indexing and Search (HANNIS) for Large Descriptor Databases

2023-01-26 · IEEE International Conference on Big Data (Big Data) 2023 1 · M M Mahabubur Rahman, Jelena Tešić

In this paper, we present a novel method for efficient and effective retrieval of similar deep descriptors. Our new hybrid method for indexing and searching for the approximate nearest neighbors in high-dimensional large…

Retrieval

Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs

2024-08-20 · Maxim Ifergan, Leshem Choshen, Roee Aharoni, Idan Szpektor 외

The veracity of a factoid is largely independent of the language it is written in. However, language models are inconsistent in their ability to answer the same factual question across languages. This raises questions ab…

knowledge editing