paper-with-me

홈 › Papers

AssistQ: Affordance-centric Question-driven Task Completion for Egocentric Assistant

2022-03-08 · Benita Wong, Joya Chen, You Wu, Stan Weixian Lei, Dongxing Mao, Difei Gao, Mike Zheng Shou

A long-standing goal of intelligent assistants such as AR glasses/robots has been to assist users in affordance-centric real-world scenarios, such as "how can I run the microwave for 1 minute?". However, there is still no clear task definition and suitable benchmarks. In this paper, we define a new task called Affordance-centric Question-driven Task Completion, where the AI assistant should learn from instructional videos to provide step-by-step help in the user's view. To support the task, we constructed AssistQ, a new dataset comprising 531 question-answer samples from 100 newly filmed instructional videos. We also developed a novel Question-to-Actions (Q2A) model to address the AQTC task and validate it on the AssistQ dataset. The results show that our model significantly outperforms several VQA-related baselines while still having large room for improvement. We expect our task and dataset to advance Egocentric AI Assistant's development. Our project page is available at: https://showlab.github.io/assistq/.

📄 PDF Abstract BibTeX arXiv:2203.04203

Code (4)

showlab/Q2A 공식 구현 pytorch
jaykim9870/cvpr-22_loveu_unipyler pytorch
starsholic/loveu-cvpr22-aqtc pytorch
zcfinal/loveu-cvpr23-aqtc pytorch

Tasks

Visual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Winning the CVPR'2022 AQTC Challenge: A Two-stage Function-centric Approach

2022-06-20 · Shiwei Wu, Weidong He, Tong Xu, Hao Wang 외

Affordance-centric Question-driven Task Completion for Egocentric Assistant(AQTC) is a novel task which helps AI assistant learn from instructional videos and scripts and guide the user step-by-step. In this paper, we de…

Grounding 3D Scene Affordance From Egocentric Interactions

2024-09-29 · Cuiyu Liu, Wei Zhai, Yuhang Yang, Hongchen Luo 외

Grounding 3D scene affordance aims to locate interactive regions in 3D environments, which is crucial for embodied agents to interact intelligently with their surroundings. Most existing approaches achieve this by mappin…

Text-driven Affordance Learning from Egocentric Vision

2024-04-03 · Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, Shinsuke Mori

Visual affordance learning is a key component for robots to understand how to interact with objects. Conventional approaches in this field rely on pre-defined objects and actions, falling short of capturing diverse inter…

Referring ExpressionReferring Expression Comprehension

Learning Affordance Grounding from Exocentric Images

2022-03-18 · CVPR 2022 1 · Hongchen Luo, Wei Zhai, Jing Zhang, Yang Cao 외

Affordance grounding, a task to ground (i.e., localize) action possibility region in objects, which faces the challenge of establishing an explicit link with object parts due to the diversity of interactive affordance. H…

DiversityHuman-Object Interaction DetectionObjectTransfer Learning

Grounded Affordance from Exocentric View

2022-08-28 · Hongchen Luo, Wei Zhai, Jing Zhang, Yang Cao 외

Affordance grounding aims to locate objects' "action possibilities" regions, which is an essential step toward embodied intelligence. Due to the diversity of interactive affordance, the uniqueness of different individual…

DiversityHuman-Object Interaction DetectionObjectTransfer Learning