paper-with-me

홈 › Papers

TEXT2TASTE: A Versatile Egocentric Vision System for Intelligent Reading Assistance Using Large Language Model

2024-04-14 · Wiktor Mucha, Florin Cuconasu, Naome A. Etori, Valia Kalokyri, Giovanni Trappolini

The ability to read, understand and find important information from written text is a critical skill in our daily lives for our independence, comfort and safety. However, a significant part of our society is affected by partial vision impairment, which leads to discomfort and dependency in daily activities. To address the limitations of this part of society, we propose an intelligent reading assistant based on smart glasses with embedded RGB cameras and a Large Language Model (LLM), whose functionality goes beyond corrective lenses. The video recorded from the egocentric perspective of a person wearing the glasses is processed to localise text information using object detection and optical character recognition methods. The LLM processes the data and allows the user to interact with the text and responds to a given query, thus extending the functionality of corrective lenses with the ability to find and summarize knowledge from the text. To evaluate our method, we create a chat-based application that allows the user to interact with the system. The evaluation is conducted in a real-world setting, such as reading menus in a restaurant, and involves four participants. The results show robust accuracy in text retrieval. The system not only provides accurate meal suggestions but also achieves high user satisfaction, highlighting the potential of smart glasses and LLMs in assisting people with special needs.

📄 PDF Abstract BibTeX arXiv:2404.09254

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingLarge Language Modelobject-detectionObject DetectionOptical Character RecognitionText Retrieval

Similar Papers 제목 키워드 기반

EgoLM: Multi-Modal Language Model of Egocentric Motions

2024-09-26 · CVPR 2025 1 · Fangzhou Hong, Vladimir Guzov, Hyo Jin Kim, Yuting Ye 외

As the prevalence of wearable devices, learning egocentric motions becomes essential to develop contextual AI. In this work, we present EgoLM, a versatile framework that tracks and understands egocentric motions from mul…

Language ModelingLanguage ModellingmodelMotion Generation+1

A Survey on Recent Advances of Computer Vision Algorithms for Egocentric Video

2015-01-12 · Sven Bambach

Recent technological advances have made lightweight, head mounted cameras both practical and affordable and products like Google Glass show first approaches to introduce the idea of egocentric (first-person) video to the…

Action DetectionActivity DetectionObject RecognitionVideo Summarization

Visual Summary of Egocentric Photostreams by Representative Keyframes

2015-05-05 · Marc Bolaños, Ricard Mestre, Estefanía Talavera, Xavier Giró-i-Nieto 외

Building a visual summary from an egocentric photostream captured by a lifelogging wearable camera is of high interest for different applications (e.g. memory reinforcement). In this paper, we propose a new summarization…

Clustering

TASTE-Streaming: Towards Streamable Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling

2026-03-12 · Liang-Hsuan Tseng, Hung-yi Lee arxiv

Text-speech joint spoken language modeling (SLM) aims at natural and intelligent speech-based interactions, but developing such a system may suffer from modality mismatch: speech unit sequences are much longer than text …

ZeroWBC: Learning Natural Whole-Body Humanoid Interaction from Human Egocentric Data

2026-03-10 · Haoran Yang, Jiacheng Bao, Yucheng Xin, Haoming Song 외 arxiv

Achieving versatile and natural whole-body humanoid interaction control remains challenging due to the high cost of whole-body teleoperation data. We present ZeroWBC, a teleoperation-free framework that learns humanoid w…