paper-with-me

홈 › Papers

Egocentric Image Captioning for Privacy-Preserved Passive Dietary Intake Monitoring

2021-07-01 · Jianing Qiu, Frank P. -W. Lo, Xiao Gu, Modou L. Jobarteh, Wenyan Jia, Tom Baranowski, Matilda Steiner-Asiedu, Alex K. Anderson, Megan A McCrory, Edward Sazonov, Mingui Sun, Gary Frost, Benny Lo

Camera-based passive dietary intake monitoring is able to continuously capture the eating episodes of a subject, recording rich visual information, such as the type and volume of food being consumed, as well as the eating behaviours of the subject. However, there currently is no method that is able to incorporate these visual clues and provide a comprehensive context of dietary intake from passive recording (e.g., is the subject sharing food with others, what food the subject is eating, and how much food is left in the bowl). On the other hand, privacy is a major concern while egocentric wearable cameras are used for capturing. In this paper, we propose a privacy-preserved secure solution (i.e., egocentric image captioning) for dietary assessment with passive monitoring, which unifies food recognition, volume estimation, and scene understanding. By converting images into rich text descriptions, nutritionists can assess individual dietary intake based on the captions instead of the original images, reducing the risk of privacy leakage from images. To this end, an egocentric dietary image captioning dataset has been built, which consists of in-the-wild images captured by head-worn and chest-worn cameras in field studies in Ghana. A novel transformer-based architecture is designed to caption egocentric dietary images. Comprehensive experiments have been conducted to evaluate the effectiveness and to justify the design of the proposed architecture for egocentric dietary image captioning. To the best of our knowledge, this is the first work that applies image captioning for dietary intake assessment in real life settings.

📄 PDF Abstract BibTeX arXiv:2107.00372

Code (0)

등록된 구현이 없습니다.

Tasks

Food RecognitionImage CaptioningScene Understanding

Similar Papers 제목 키워드 기반

Retrieval-Augmented Egocentric Video Captioning

2024-01-01 · CVPR 2024 1 · Jilan Xu, Yifei HUANG, Junlin Hou, Guo Chen 외

Understanding human actions from videos of first-person view poses significant challenges. Most prior approaches explore representation learning on egocentric videos only, while overlooking the potential benefit of explo…

Representation LearningRetrievalVideo Captioning

Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos

2023-11-28 · Takehiko Ohkawa, Takuma Yagi, Taichi Nishimura, Ryosuke Furuta 외

We propose a novel benchmark for cross-view knowledge transfer of dense video captioning, adapting models from web instructional videos with exocentric views to an egocentric view. While dense video captioning (predictin…

Dense Video CaptioningTransfer LearningVideo Captioning

Anonymizing Egocentric Videos

2021-01-01 · ICCV 2021 10 · Daksh Thapar, Aditya Nigam, Chetan Arora

In egocentric videos, the face of a wearer capturing the video is never captured. This gives a false sense of security that the wearer's privacy is preserved while sharing such videos. However, egocentric cameras are…

Activity Recognitionobject-detectionObject DetectionOptical Flow Estimation

Sensor-Augmented Egocentric-Video Captioning with Dynamic Modal Attention

2021-09-07 · Katsuyuki Nakamura, Hiroki Ohashi, Mitsuhiro Okada

Automatically describing video, or video captioning, has been widely studied in the multimedia field. This paper proposes a new task of sensor-augmented egocentric-video captioning, a newly constructed dataset for it cal…

Sensor FusionVideo Captioning

Spontaneous Spatial Cognition Emerges during Egocentric Video Viewing through Non-invasive BCI

2025-07-16 · Weichen Dai, Yuxuan Huang, Li Zhu, Dongjun Liu 외 arxiv

Humans possess a remarkable capacity for spatial cognition, allowing for self-localization even in novel or unfamiliar environments. While hippocampal neurons encoding position and orientation are well documented, the la…