paper-with-me

홈 › Papers

MET-Bench: Multimodal Entity Tracking for Evaluating the Limitations of Vision-Language and Reasoning Models

2025-02-15 · Vanya Cohen, Raymond Mooney

Entity tracking is a fundamental challenge in natural language understanding, requiring models to maintain coherent representations of entities. Previous work has benchmarked entity tracking performance in purely text-based tasks. We introduce MET-Bench, a multimodal entity tracking benchmark designed to evaluate the ability of vision-language models to track entity states across modalities. Using two structured domains, Chess and the Shell Game, we assess how effectively current models integrate textual and image-based state updates. Our findings reveal a significant performance gap between text-based and image-based tracking and that this performance gap stems from deficits in visual reasoning rather than perception. We further show that explicit text-based reasoning strategies improve performance, yet substantial limitations remain, especially in long-horizon multimodal scenarios. Our results highlight the need for improved multimodal representations and reasoning techniques to bridge the gap between textual and visual entity tracking.

📄 PDF Abstract BibTeX arXiv:2502.10886

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language UnderstandingVisual Reasoning

Similar Papers 제목 키워드 기반

WaspMOT: A Benchmark for Long-Term Multi-Object Tracking of Trichogramma Wasps

2026-07-09 · Tomasz Stanczyk, Yuan Gao, Hardik Agarwal, Seongroo Yoon 외 arxiv

Multi-object tracking (MOT) has achieved strong performance on benchmarks dominated by short video sequences. However, such datasets do not adequately evaluate long-term identity preservation, where objects must be track…

Multi-Object Tracking

MMKE-Bench: A Multimodal Editing Benchmark for Diverse Visual Knowledge

2025-02-27 · Yuntao Du, Kailin Jiang, Zhi Gao, Chenrui Shi 외

Knowledge editing techniques have emerged as essential tools for updating the factual knowledge of large language models (LLMs) and multimodal models (LMMs), allowing them to correct outdated or inaccurate information wi…

knowledge editing

FEMOT: Multi-Object Tracking using Frame and Event Cameras

2026-06-12 · Shiao Wang, Xiao Wang, Chao Wang, Yitao Li 외 arxiv

Conventional RGB cameras have been widely used in multi-object tracking due to their ability to capture rich appearance and semantic information. However, their performance is often degraded under complex real-world chal…

Multi-Object TrackingObject Localization

ESBM: An Entity Summarization BenchMark

2020-03-08 · Qingxia Liu, Gong Cheng, Kalpa Gunaratna, Yuzhong Qu

Entity summarization is the problem of computing an optimal compact summary for an entity by selecting a size-constrained subset of triples from RDF data. Entity summarization supports a multiplicity of applications and …

ChartEditBench: Evaluating Grounded Multi-Turn Chart Editing in Multimodal Language Models

2026-02-17 · Manav Nitin Kapadnis, Lawanya Baghel, Atharva Naik, Carolyn Rosé arxiv

While Multimodal Large Language Models (MLLMs) perform strongly on single-turn chart generation, their ability to support real-world exploratory data analysis remains underexplored. In practice, users iteratively refine …