paper-with-me

Papers

Predicting the Next Action by Modeling the Abstract Goal

2022-09-12 · Debaditya Roy, Basura Fernando

The problem of anticipating human actions is an inherently uncertain one. However, we can reduce this uncertainty if we have a sense of the goal that the actor is trying to achieve. Here, we present an action anticipation model that leverages goal information for the purpose of reducing the uncertainty in future predictions. Since we do not possess goal information or the observed actions during inference, we resort to visual representation to encapsulate information about both actions and goals. Through this, we derive a novel concept called abstract goal which is conditioned on observed sequences of visual features for action anticipation. We design the abstract goal as a distribution whose parameters are estimated using a variational recurrent network. We sample multiple candidates for the next action and introduce a goal consistency measure to determine the best candidate that follows from the abstract goal. Our method obtains impressive results on the very challenging Epic-Kitchens55 (EK55), EK100, and EGTEA Gaze+ datasets. We obtain absolute improvements of +13.69, +11.24, and +5.19 for Top-1 verb, Top-1 noun, and Top-1 action anticipation accuracy respectively over prior state-of-the-art methods for seen kitchens (S1) of EK55. Similarly, we also obtain significant improvements in the unseen kitchens (S2) set for Top-1 verb (+10.75), noun (+5.84) and action (+2.87) anticipation. Similar trend is observed for EGTEA Gaze+ dataset, where absolute improvement of +9.9, +13.1 and +6.8 is obtained for noun, verb, and action anticipation. It is through the submission of this paper that our method is currently the new state-of-the-art for action anticipation in EK55 and EGTEA Gaze+ https://competitions.codalab.org/competitions/20071#results Code available at https://github.com/debadityaroy/Abstract_Goal

📄 PDF Abstract BibTeX arXiv:2209.05044

Code (0)

등록된 구현이 없습니다.

Tasks

Action Anticipation

Similar Papers 제목 키워드 기반

Machine Learning for Performance Prediction of Channel Bonding in Next-Generation IEEE 802.11 WLANs

2021-05-29 · Francesc Wilhelmi, David Góez, Paola Soto, Ramon Vallés 외

With the advent of Artificial Intelligence (AI)-empowered communications, industry, academia, and standardization organizations are progressing on the definition of mechanisms and procedures to address the increasing com…

BIG-bench Machine Learning

InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision

2025-12-01 · Chenting Wang, Yuhan Zhu, Yicheng Xu, Jiange Yang 외 arxiv

Large-scale video-text pretraining achieves strong performance but depends on noisy, synthetic captions with limited semantic coverage, often overlooking implicit world knowledge such as object motion, 3D geometry, and p…

Representation Learning

Real-World Robot Control by Deep Active Inference With a Temporally Hierarchical World Model

2025-12-01 · Kentaro Fujii, Shingo Murata arxiv

Robots in uncertain real-world environments must perform both goal-directed and exploratory actions. However, most deep learning-based control methods neglect exploration and struggle under uncertainty. To address this, …

Offline Policy Learning via Skill-step Abstraction for Long-horizon Goal-Conditioned Tasks

2024-08-21 · Donghoon Kim, Minjong Yoo, Honguk Woo

Goal-conditioned (GC) policy learning often faces a challenge arising from the sparsity of rewards, when confronting long-horizon goals. To address the challenge, we explore skill-based GC policy learning in offline sett…

parameter-efficient fine-tuning

End-to-End Personalized Next Location Recommendation via Contrastive User Preference Modeling

2023-03-22 · Yan Luo, Ye Liu, Fu-Lai Chung, Yu Liu 외

Predicting the next location is a highly valuable and common need in many location-based services such as destination prediction and route planning. The goal of next location recommendation is to predict the next point-o…

Decoder