paper-with-me

홈 › Papers

Winning the CVPR'2022 AQTC Challenge: A Two-stage Function-centric Approach

2022-06-20 · Shiwei Wu, Weidong He, Tong Xu, Hao Wang, Enhong Chen

Affordance-centric Question-driven Task Completion for Egocentric Assistant(AQTC) is a novel task which helps AI assistant learn from instructional videos and scripts and guide the user step-by-step. In this paper, we deal with the AQTC via a two-stage Function-centric approach, which consists of Question2Function Module to ground the question with the related function and Function2Answer Module to predict the action based on the historical steps. We evaluated several possible solutions in each module and obtained significant gains compared to the given baselines. Our code is available at \url{https://github.com/starsholic/LOVEU-CVPR22-AQTC}.

📄 PDF Abstract BibTeX arXiv:2206.09597

Code (1)

starsholic/loveu-cvpr22-aqtc 공식 구현 pytorch

Similar Papers 제목 키워드 기반

Technical Report for CVPR 2022 LOVEU AQTC Challenge

2022-06-29 · Hyeonyu Kim, Jongeun Kim, Jeonghun Kang, Sanguk Park 외

This technical report presents the 2nd winning model for AQTC, a task newly introduced in CVPR 2022 LOng-form VidEo Understanding (LOVEU) challenges. This challenge faces difficulties with multi-step answers, multi-modal…

Video Understanding

First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment

2023-06-23 · Tom Tongjia Chen, Hongshan Yu, Zhengeng Yang, Ming Li 외

Affordance-Centric Question-driven Task Completion (AQTC) has been proposed to acquire knowledge from videos to furnish users with comprehensive and systematic instructions. However, existing methods have hitherto neglec…

Human-Object Interaction Detection

A Solution to CVPR'2023 AQTC Challenge: Video Alignment for Multi-Step Inference

2023-06-26 · Chao Zhang, Shiwei Wu, Sirui Zhao, Tong Xu 외

Affordance-centric Question-driven Task Completion (AQTC) for Egocentric Assistant introduces a groundbreaking scenario. In this scenario, through learning instructional videos, AI assistants provide users with step-by-s…

Video Alignment

Geospatial Foundational Embedder: Top-1 Winning Solution on EarthVision Embed2Scale Challenge (CVPR 2025)

2025-09-03 · Zirui Xu, Raphael Tang, Mike Bianco, Qi Zhang 외 arxiv

EarthVision Embed2Scale challenge (CVPR 2025) aims to develop foundational geospatial models to embed SSL4EO-S12 hyperspectral geospatial data cubes into embedding vectors that faciliatetes various downstream tasks, e.g.…

Real-Time Anchor-Free Single-Stage 3D Detection with IoU-Awareness

2021-07-29 · Runzhou Ge, Zhuangzhuang Ding, Yihan Hu, Wenxin Shao 외

In this report, we introduce our winning solution to the Real-time 3D Detection and also the "Most Efficient Model" in the Waymo Open Dataset Challenges at CVPR 2021. Extended from our last year's award-winning model AFD…

Data AugmentationGPU