paper-with-me

홈 › Papers

Lost in Translation and Noise: A Deep Dive into the Failure Modes of VLMs on Real-World Tables

2025-11-21 · Anshul Singh, Rohan Chaudhary, Gagneet Singh, Abhay Kumary arxiv

The impressive performance of VLMs is largely measured on benchmarks that fail to capture the complexities of real-world scenarios. Existing datasets for tabular QA, such as WikiTableQuestions and FinQA, are overwhelmingly monolingual (English) and present tables in a digitally perfect, clean format. This creates a significant gap between research and practice. To address this, we present \textbf{MirageTVQA}, a new benchmark designed to evaluate VLMs on these exact dimensions. Featuring nearly 60,000 QA pairs across 24 languages, MirageTVQA challenges models with tables that are not only multilingual but also visually imperfect, incorporating realistic noise to mimic scanned documents. Our evaluation of the leading VLMs reveals two primary failure points: a severe degradation in performance (over 35\% drop for the best models) when faced with visual noise and a consistent English-first bias where reasoning abilities fail to transfer to other languages. MirageTVQA provides a benchmark for measuring and driving progress towards more robust VLM models for table reasoning. The dataset and the code are available at: https://github.com/anshulsc/MirageTVQA.

📄 PDF Abstract BibTeX arXiv:2511.17238

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Lost in Back-Translation: Emotion Preservation in Neural Machine Translation

2020-12-01 · COLING 2020 8 · Enrica Troiano, Roman Klinger, Sebastian Pad{\'o}

Machine translation provides powerful methods to convert text between languages, and is therefore a technology enabling a multilingual world. An important part of communication, however, takes place at the non-propositio…

DiversityMachine TranslationRe-RankingStyle Transfer+1

Towards Tailored Recovery of Lexical Diversity in Literary Machine Translation

2024-08-30 · Esther Ploeger, Huiyuan Lai, Rik van Noord, Antonio Toral

Machine translations are found to be lexically poorer than human translations. The loss of lexical diversity through MT poses an issue in the automatic translation of literature, where it matters not only what is written…

DiversityMachine TranslationRerankingTranslation

Lost in Translation: Loss and Decay of Linguistic Richness in Machine Translation

2019-06-28 · WS 2019 8 · Eva Vanmassenhove, Dimitar Shterionov, Andy Way

This work presents an empirical approach to quantifying the loss of lexical richness in Machine Translation (MT) systems compared to Human Translation (HT). Our experiments show how current MT systems indeed fail to rend…

DiversityMachine TranslationTranslation

MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation

2025-10-28 · Parker Riley, Daniel Deutsch, Mara Finkelstein, Colten DiIanni 외 arxiv

Human evaluation of machine translation is in an arms race with translation model quality: as our models get better, our evaluation methods need to be improved to ensure that quality gains are not lost in evaluation nois…

Machine Translation

Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss

2026-04-14 · Ronald Skorobogat, Ameya Prabhu, Matthias Bethge arxiv

Multilingual benchmarks guide the development of frontier models. Yet multilingual evaluations reported by frontier models are structured similar to popular reasoning and knowledge benchmarks, but across many languages. …

Mathematical Reasoning