GLIMPSE : Real-Time Text Recognition and Contextual Understanding for VQA in Wearables
Video Large Language Models (Video LLMs) have shown remarkable progress in understanding and reasoning about visual content, particularly in tasks involving text recognition and text-based visual question answering (Text VQA). However, deploying Text VQA on wearable devices faces a fundamental tension: text recognition requires high-resolution video, but streaming high-quality video drains battery and causes thermal throttling. Moreover, existing models struggle to maintain coherent temporal context when processing text across multiple frames in real-time streams. We observe that text recognition and visual reasoning have asymmetric resolution requirements - OCR needs fine detail while scene understanding tolerates coarse features. We exploit this asymmetry with a hybrid architecture that performs selective high-resolution OCR on-device while streaming low-resolution video for visual context. On a benchmark of text-based VQA samples across five task categories, our system achieves 72% accuracy at 0.49x the power consumption of full-resolution streaming, enabling sustained VQA sessions on resource-constrained wearables without sacrificing text understanding quality.
Code (0)
등록된 구현이 없습니다.
Tasks
Visual Question AnsweringScene UnderstandingVisual ReasoningSimilar Papers 제목 키워드 기반
Glimpse Clouds: Human Activity Recognition from Unstructured Feature Points
We propose a method for human activity recognition from RGB data that does not rely on any pose information during test time and does not explicitly calculate pose information internally. Instead, a visual attention modu…
Action RecognitionActivity PredictionActivity RecognitionHuman Activity Recognition+2RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition
The attention-based encoder-decoder framework has recently achieved impressive results for scene text recognition, and many variants have emerged with improvements in recognition quality. However, it performs poorly on c…
DecoderIrregular Text RecognitionPositionScene Text RecognitionMulti-Glimpse LSTM with Color-Depth Feature Fusion for Human Detection
With the development of depth cameras such as Kinect and Intel Realsense, RGB-D based human detection receives continuous research attention due to its usage in a variety of applications. In this paper, we propose a new …
Human DetectionRecurrent Attention Models with Object-centric Capsule Representation for Multi-object Recognition
The visual system processes a scene using a sequence of selective glimpses, each driven by spatial and object-based attention. These glimpses reflect what is relevant to the ongoing task and are selected through recurren…
DecoderObjectObject RecognitionGliTr: Glimpse Transformers with Spatiotemporal Consistency for Online Action Prediction
Many online action prediction models observe complete frames to locate and attend to informative subregions in the frames called glimpses and recognize an ongoing action based on global and local information. However, in…
Action Recognition