paper-with-me

홈 › Papers

First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment

2023-06-23 · Tom Tongjia Chen, Hongshan Yu, Zhengeng Yang, Ming Li, Zechuan Li, Jingwen Wang, Wei Miao, Wei Sun, Chen Chen

Affordance-Centric Question-driven Task Completion (AQTC) has been proposed to acquire knowledge from videos to furnish users with comprehensive and systematic instructions. However, existing methods have hitherto neglected the necessity of aligning spatiotemporal visual and linguistic signals, as well as the crucial interactional information between humans and objects. To tackle these limitations, we propose to combine large-scale pre-trained vision-language and video-language models, which serve to contribute stable and reliable multimodal data and facilitate effective spatiotemporal visual-textual alignment. Additionally, a novel hand-object-interaction (HOI) aggregation module is proposed which aids in capturing human-object interaction information, thereby further augmenting the capacity to understand the presented scenario. Our method achieved first place in the CVPR'2023 AQTC Challenge, with a Recall@1 score of 78.7\%. The code is available at https://github.com/tomchen-ctj/CVPR23-LOVEU-AQTC.

📄 PDF Abstract BibTeX arXiv:2306.13380

Code (1)

tomchen-ctj/cvpr23-loveu-aqtc 공식 구현 pytorch

Tasks

Human-Object Interaction Detection

Similar Papers 제목 키워드 기반

A Solution to CVPR'2023 AQTC Challenge: Video Alignment for Multi-Step Inference

2023-06-26 · Chao Zhang, Shiwei Wu, Sirui Zhao, Tong Xu 외

Affordance-centric Question-driven Task Completion (AQTC) for Egocentric Assistant introduces a groundbreaking scenario. In this scenario, through learning instructional videos, AI assistants provide users with step-by-s…

Video Alignment

Technical Report for CVPR 2022 LOVEU AQTC Challenge

2022-06-29 · Hyeonyu Kim, Jongeun Kim, Jeonghun Kang, Sanguk Park 외

This technical report presents the 2nd winning model for AQTC, a task newly introduced in CVPR 2022 LOng-form VidEo Understanding (LOVEU) challenges. This challenge faces difficulties with multi-step answers, multi-modal…

Video Understanding

Winning the CVPR'2022 AQTC Challenge: A Two-stage Function-centric Approach

2022-06-20 · Shiwei Wu, Weidong He, Tong Xu, Hao Wang 외

Affordance-centric Question-driven Task Completion for Egocentric Assistant(AQTC) is a novel task which helps AI assistant learn from instructional videos and scripts and guide the user step-by-step. In this paper, we de…

The Second-place Solution for CVPR 2022 SoccerNet Tracking Challenge

2022-11-24 · Fan Yang, Shigeyuki Odashima, Shoichi Masui, Shan Jiang

This is our second-place solution for CVPR 2022 SoccerNet Tracking Challenge. Our method mainly includes two steps: online short-term tracking using our Cascaded Buffer-IoU (C-BIoU) Tracker, and, offline long-term tracki…

Clustering

The Third Place Solution for CVPR2022 AVA Accessibility Vision and Autonomy Challenge

2022-06-28 · Bo Yan, Leilei Cao, Zhuang Li, Hongbin Wang

The goal of AVA challenge is to provide vision-based benchmarks and methods relevant to accessibility. In this paper, we introduce the technical details of our submission to the CVPR2022 AVA Challenge. Firstly, we conduc…

Data Augmentation