Lost in Motion: Vision Language Models Fail the Dynamic Gauges Test
The digital transformation of industrial manufacturing increasingly relies on the ability of autonomous robots to interact with legacy infrastructure, particularly analog gauges. Vision-Language Models (VLMs) have the potential to provide a general solution for gauge reading and have already shown good performance in instrument recognition. However, performing accurate, real-time gauge readings is a more complex task. This paper evaluates state-of-the-art models, including versions from the GPT-5 (5.4 Thinking and 5.3 Instant) and Gemini 3 (Pro and Flash) families, against a set of simple but realistic dynamic gauge reading scenarios. To facilitate this evaluation, we introduce a novel dataset comprising video sequences of three instruments of different gauge types: circular, linear, and Vernier, under diverse motion and speed profiles. Our findings indicate that the evaluated frontier VLMs, under our specific testing conditions, exhibit a limited ability to interpret needle trajectories and scale semantics, failing to provide the traceability and reliability needed for safety-critical monitoring. The results demonstrate that these models have not yet achieved the performance necessary to be classified as trustworthy synthetic instruments under existing IEEE and ISO standards.
Code (0)
등록된 구현이 없습니다.
Tasks
Instrument RecognitionSimilar Papers 제목 키워드 기반
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
We study whether vision-language models (VLMs) can solve relative camera pose estimation (RCPE) from image pairs, a direct test of multi-view spatial reasoning. We cast RCPE as a discrete verbal classification task and i…
Camera Pose EstimationSpatial ReasoningSpatio-Temporal Difference Guided Motion Deblurring with the Complementary Vision Sensor
Motion blur arises when rapid scene changes occur during the exposure period, collapsing rich intra-exposure motion into a single RGB frame. Without explicit structural or temporal cues, RGB-only deblurring is highly ill…
Vision-Motion-Reference Alignment for Referring Multi-Object Tracking via Multi-Modal Large Language Models
Referring Multi-Object Tracking (RMOT) extends conventional multi-object tracking (MOT) by introducing natural language references for multi-modal fusion tracking. RMOT benchmarks only describe the object's appearance, r…
Multi-Object TrackingCL4D: Contrastive Language-4D Pretraining for Vision-Language Reasoning in Dynamic Scenes
4D understanding and reasoning is a fundamental capability for embodied AI agents operating in dynamic physical environments. However, existing vision encoders are largely limited to static 2D images or 3D point clouds w…
Contrastive LearningPoint CloudsLost in Back-Translation: Emotion Preservation in Neural Machine Translation
Machine translation provides powerful methods to convert text between languages, and is therefore a technology enabling a multilingual world. An important part of communication, however, takes place at the non-propositio…
DiversityMachine TranslationRe-RankingStyle Transfer+1