paper-with-me

Papers

Knowledge Guided Learning: Towards Open Domain Egocentric Action Recognition with Zero Supervision

2020-09-16 · Sathyanarayanan N. Aakur, Sanjoy Kundu, Nikhil Gunti

Advances in deep learning have enabled the development of models that have exhibited a remarkable tendency to recognize and even localize actions in videos. However, they tend to experience errors when faced with scenes or examples beyond their initial training environment. Hence, they fail to adapt to new domains without significant retraining with large amounts of annotated data. In this paper, we propose to overcome these limitations by moving to an open-world setting by decoupling the ideas of recognition and reasoning. Building upon the compositional representation offered by Grenander's Pattern Theory formalism, we show that attention and commonsense knowledge can be used to enable the self-supervised discovery of novel actions in egocentric videos in an open-world setting, where data from the observed environment (the target domain) is open i.e., the vocabulary is partially known and training examples (both labeled and unlabeled) are not available. We show that our approach can infer and learn novel classes for open vocabulary classification in egocentric videos and novel object detection with zero supervision. Extensive experiments show its competitive performance on two publicly available egocentric action recognition datasets (GTEA Gaze and GTEA Gaze+) under open-world conditions.

📄 PDF Abstract BibTeX arXiv:2009.07470

Code (0)

등록된 구현이 없습니다.

Tasks

Action RecognitionDomain AdaptationNovel Object Detectionobject-detectionObject DetectionZero-Shot Learning

Similar Papers 제목 키워드 기반

From My View to Yours: Ego-Augmented Learning in Large Vision Language Models for Understanding Exocentric Daily Living Activities

2025-01-10 · Dominick Reilly, Manish Kumar Govind, Le Xue, Srijan Das

Large Vision Language Models (LVLMs) have demonstrated impressive capabilities in video understanding, yet their adoption for Activities of Daily Living (ADL) remains limited by their inability to capture fine-grained in…

Human-Object Interaction DetectionKnowledge DistillationVideo Understanding

EgoX: Egocentric Video Generation from a Single Exocentric Video

2025-12-09 · Taewoong Kang, Kinam Kim, Dohyeon Kim, Minho Park 외 arxiv

Egocentric perception enables humans to experience and understand the world directly from their own point of view. Translating exocentric (third-person) videos into egocentric (first-person) videos opens up new possibili…

Video Generation

Closed-Loop Transfer for Weakly-supervised Affordance Grounding

2025-10-20 · Jiajin Tang, Zhengxuan Wei, Ge Zheng, Sibei Yang arxiv

Humans can perform previously unexperienced interactions with novel objects simply by observing others engage with them. Weakly-supervised affordance grounding mimics this process by learning to locate object regions tha…

Knowledge Distillation

Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding

2025-03-12 · Haoyu Zhang, Qiaohui Chu, Meng Liu, Yunxiao Wang 외

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. Current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exoce…

Instruction FollowingVideo Understanding

EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing

2025-12-05 · Runjia Li, Moayed Haji-Ali, Ashkan Mirzaei, Chaoyang Wang 외 arxiv

We study instruction-guided editing of egocentric videos for interactive AR applications. While recent AI video editors perform well on third-person footage, egocentric views present unique challenges - including rapid e…