paper-with-me

Papers

PAL: Intelligence Augmentation using Egocentric Visual Context Detection

2021-05-22 · Mina Khan, Pattie Maes

Egocentric visual context detection can support intelligence augmentation applications. We created a wearable system, called PAL, for wearable, personalized, and privacy-preserving egocentric visual context detection. PAL has a wearable device with a camera, heart-rate sensor, on-device deep learning, and audio input/output. PAL also has a mobile/web application for personalized context labeling. We used on-device deep learning models for generic object and face detection, low-shot custom face and context recognition (e.g., activities like brushing teeth), and custom context clustering (e.g., indoor locations). The models had over 80\% accuracy in in-the-wild contexts (~1000 images) and we tested PAL for intelligence augmentation applications like behavior change. We have made PAL is open-source to further support intelligence augmentation using personalized and privacy-preserving egocentric visual contexts.

📄 PDF Abstract BibTeX arXiv:2105.10735

Code (0)

등록된 구현이 없습니다.

Tasks

ClusteringDeep LearningFace DetectionPrivacy Preserving

Similar Papers 제목 키워드 기반

HERO-VQL: Hierarchical, Egocentric and Robust Visual Query Localization

2025-08-30 · Joohyun Chang, Soyeon Hong, Hyogun Lee, Seong Jong Ha 외 arxiv

In this work, we tackle the egocentric visual query localization (VQL), where a model should localize the query object in a long-form egocentric video. Frequent and abrupt viewpoint changes in egocentric videos cause sig…

Object LocalizationObject Recognition

Reinforcing Egocentric Spatial Perception in Multimodal Large Language Models via Ego Scene Augmentation

2026-07-16 · Chi Kit Wong, Ye Pan, Yuanhuiyi Lyu, Xu Zheng 외 arxiv

Egocentric Visual Question Answering (VQA) has attracted widespread attention as an important task for enabling Multimodal Large Language Models (MLLMs) to interact with the real world. However, existing MLLMs struggle t…

Visual Question AnsweringSpatial Reasoning

Where is my Wallet? Modeling Object Proposal Sets for Egocentric Visual Query Localization

2022-11-18 · CVPR 2023 1 · Mengmeng Xu, Yanghao Li, Cheng-Yang Fu, Bernard Ghanem 외

This paper deals with the problem of localizing objects in image and video datasets from visual exemplars. In particular, we focus on the challenging problem of egocentric visual query localization. We first identify gra…

Object

Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration

2025-02-20 · Pengxiang Ding, Jianfei Ma, Xinyang Tong, Binghong Zou 외

This paper addresses the limitations of current humanoid robot control frameworks, which primarily rely on reactive mechanisms and lack autonomous interaction capabilities due to data scarcity. We propose Humanoid-VLA, a…

Data AugmentationHumanoid ControlMotion Generation

EgoAVU: Egocentric Audio-Visual Understanding

2026-02-05 · Ashish Seth, Xinhao Mei, Changsheng Zhao, Varun Nagaraja 외 arxiv

Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labe…