paper-with-me

홈 › Papers

EgoQR: Efficient QR Code Reading in Egocentric Settings

2024-10-07 · Mohsen Moslehpour, Yichao Lu, Pierce Chuang, Ashish Shenoy, Debojeet Chatterjee, Abhay Harpale, Srihari Jayakumar, Vikas Bhardwaj, Seonghyeon Nam, Anuj Kumar

QR codes have become ubiquitous in daily life, enabling rapid information exchange. With the increasing adoption of smart wearable devices, there is a need for efficient, and friction-less QR code reading capabilities from Egocentric point-of-views. However, adapting existing phone-based QR code readers to egocentric images poses significant challenges. Code reading from egocentric images bring unique challenges such as wide field-of-view, code distortion and lack of visual feedback as compared to phones where users can adjust the position and framing. Furthermore, wearable devices impose constraints on resources like compute, power and memory. To address these challenges, we present EgoQR, a novel system for reading QR codes from egocentric images, and is well suited for deployment on wearable devices. Our approach consists of two primary components: detection and decoding, designed to operate on high-resolution images on the device with minimal power consumption and added latency. The detection component efficiently locates potential QR codes within the image, while our enhanced decoding component extracts and interprets the encoded information. We incorporate innovative techniques to handle the specific challenges of egocentric imagery, such as varying perspectives, wider field of view, and motion blur. We evaluate our approach on a dataset of egocentric images, demonstrating 34% improvement in reading the code compared to an existing state of the art QR code readers.

📄 PDF Abstract BibTeX arXiv:2410.05497

Code (0)

등록된 구현이 없습니다.

Tasks

Friction

Similar Papers 제목 키워드 기반

Reading Recognition in the Wild

2025-05-30 · Charig Yang, Samiul Alam, Shakhrul Iman Siam, Michael J. Proulx 외

To enable egocentric contextual AI in always-on smart glasses, it is crucial to be able to keep a record of the user's interactions with the world, including during reading. In this paper, we introduce a new task of read…

Diversity

EgoDistill: Egocentric Head Motion Distillation for Efficient Video Understanding

2023-01-05 · NeurIPS 2023 11

Recent advances in egocentric video understanding models are promising, but their heavy computational expense is a barrier for many real-world applications. To address this challenge, we propose EgoDistill, a distillatio…

Video Understanding

Ego-Object Discovery

2015-04-07 · Marc Bolaños, Petia Radeva

Lifelogging devices are spreading faster everyday. This growth can represent great benefits to develop methods for extraction of meaningful information about the user wearing the device and his/her environment. In this p…

ObjectObject Discovery

Are We There Yet? Exploring the Capabilities of MLLMs in Assistive AI Applications

2026-06-23 · Shayon Dasgupta, Avijit Dasgupta, C. V. Jawahar arxiv

Multimodal Large Language Models (MLLMs) have redefined visual understanding by combining vision encoders with large-scale language models. This unified architecture enables strong performance on tasks like image caption…

Visual Question AnsweringImage Captioning

TEXT2TASTE: A Versatile Egocentric Vision System for Intelligent Reading Assistance Using Large Language Model

2024-04-14 · Wiktor Mucha, Florin Cuconasu, Naome A. Etori, Valia Kalokyri 외

The ability to read, understand and find important information from written text is a critical skill in our daily lives for our independence, comfort and safety. However, a significant part of our society is affected by …

Language ModelingLanguage ModellingLarge Language Modelobject-detection+3