paper-with-me

Papers

Lost in Motion: Vision Language Models Fail the Dynamic Gauges Test

2026-04-19 · Tairan Fu, Francisco Javier Santos-Martín, Javier Conde, Pedro Reviriego, Elena Merino-Gómez arxiv

The digital transformation of industrial manufacturing increasingly relies on the ability of autonomous robots to interact with legacy infrastructure, particularly analog gauges. Vision-Language Models (VLMs) have the potential to provide a general solution for gauge reading and have already shown good performance in instrument recognition. However, performing accurate, real-time gauge readings is a more complex task. This paper evaluates state-of-the-art models, including versions from the GPT-5 (5.4 Thinking and 5.3 Instant) and Gemini 3 (Pro and Flash) families, against a set of simple but realistic dynamic gauge reading scenarios. To facilitate this evaluation, we introduce a novel dataset comprising video sequences of three instruments of different gauge types: circular, linear, and Vernier, under diverse motion and speed profiles. Our findings indicate that the evaluated frontier VLMs, under our specific testing conditions, exhibit a limited ability to interpret needle trajectories and scale semantics, failing to provide the traceability and reliability needed for safety-critical monitoring. The results demonstrate that these models have not yet achieved the performance necessary to be classified as trustworthy synthetic instruments under existing IEEE and ISO standards.

📄 PDF Abstract BibTeX arXiv:2604.22829

Code (0)

등록된 구현이 없습니다.

Tasks

Instrument Recognition

Similar Papers 제목 키워드 기반

Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation

2026-01-29 · Ken Deng, Yifu Qiu, Yoni Kasten, Shay B. Cohen 외 arxiv

We study whether vision-language models (VLMs) can solve relative camera pose estimation (RCPE) from image pairs, a direct test of multi-view spatial reasoning. We cast RCPE as a discrete verbal classification task and i…

Camera Pose EstimationSpatial Reasoning

Spatio-Temporal Difference Guided Motion Deblurring with the Complementary Vision Sensor

2026-04-12 · Yapeng Meng, Lin Yang, Yuguo Chen, Xiangru Chen 외 arxiv

Motion blur arises when rapid scene changes occur during the exposure period, collapsing rich intra-exposure motion into a single RGB frame. Without explicit structural or temporal cues, RGB-only deblurring is highly ill…

Vision-Motion-Reference Alignment for Referring Multi-Object Tracking via Multi-Modal Large Language Models

2025-11-21 · Weiyi Lv, Ning Zhang, Hanyang Sun, Haoran Jiang 외 arxiv

Referring Multi-Object Tracking (RMOT) extends conventional multi-object tracking (MOT) by introducing natural language references for multi-modal fusion tracking. RMOT benchmarks only describe the object's appearance, r…

Multi-Object Tracking

CL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes

2026-08-19 · Kumal Hewagamage, Isuranga Senavirathne, Sasika Amarasinghe, Hasitha Gallella 외 arxiv

4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vision encoders are largely limited to static 2D images or 3D point clouds w…

Contrastive LearningPoint Clouds

Lost in Back-Translation: Emotion Preservation in Neural Machine Translation

2020-12-01 · COLING 2020 8 · Enrica Troiano, Roman Klinger, Sebastian Pad{\'o}

Machine translation provides powerful methods to convert text between languages, and is therefore a technology enabling a multilingual world. An important part of communication, however, takes place at the non-propositio…

DiversityMachine TranslationRe-RankingStyle Transfer+1