paper-with-me

홈 › Papers

Technical Report for CVPR 2022 LOVEU AQTC Challenge

2022-06-29 · Hyeonyu Kim, Jongeun Kim, Jeonghun Kang, Sanguk Park, Dongchan Park, Taehwan Kim

This technical report presents the 2nd winning model for AQTC, a task newly introduced in CVPR 2022 LOng-form VidEo Understanding (LOVEU) challenges. This challenge faces difficulties with multi-step answers, multi-modal, and diverse and changing button representations in video. We address this problem by proposing a new context ground module attention mechanism for more effective feature mapping. In addition, we also perform the analysis over the number of buttons and ablation study of different step networks and video features. As a result, we achieved the overall 2nd place in LOVEU competition track 3, specifically the 1st place in two out of four evaluation metrics. Our code is available at https://github.com/jaykim9870/ CVPR-22_LOVEU_unipyler.

📄 PDF Abstract BibTeX arXiv:2206.14555

Code (1)

jaykim9870/cvpr-22_loveu_unipyler 공식 구현 pytorch

Tasks

Video Understanding

Similar Papers 제목 키워드 기반

Winning the CVPR'2022 AQTC Challenge: A Two-stage Function-centric Approach

2022-06-20 · Shiwei Wu, Weidong He, Tong Xu, Hao Wang 외

Affordance-centric Question-driven Task Completion for Egocentric Assistant(AQTC) is a novel task which helps AI assistant learn from instructional videos and scripts and guide the user step-by-step. In this paper, we de…

A Solution to CVPR'2023 AQTC Challenge: Video Alignment for Multi-Step Inference

2023-06-26 · Chao Zhang, Shiwei Wu, Sirui Zhao, Tong Xu 외

Affordance-centric Question-driven Task Completion (AQTC) for Egocentric Assistant introduces a groundbreaking scenario. In this scenario, through learning instructional videos, AI assistants provide users with step-by-s…

Video Alignment

First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment

2023-06-23 · Tom Tongjia Chen, Hongshan Yu, Zhengeng Yang, Ming Li 외

Affordance-Centric Question-driven Task Completion (AQTC) has been proposed to acquire knowledge from videos to furnish users with comprehensive and systematic instructions. However, existing methods have hitherto neglec…

Human-Object Interaction Detection

Technical Report for CVPR 2024 WeatherProof Dataset Challenge: Semantic Segmentation on Paired Real Data

2024-06-09 · Guojin Cao, Jiaxu Li, Jia He, Ying Min 외

This technical report presents the implementation details of 2nd winning for CVPR'24 UG2 WeatherProof Dataset Challenge. This challenge aims at semantic segmentation of images degraded by various degrees of weather from …

Semantic Segmentation

Winning the CVPR'2021 Kinetics-GEBD Challenge: Contrastive Learning Approach

2021-06-22 · Hyolim Kang, Jinwoo Kim, KyungMin Kim, Taehyun Kim 외

Generic Event Boundary Detection (GEBD) is a newly introduced task that aims to detect "general" event boundaries that correspond to natural human perception. In this paper, we introduce a novel contrastive learning base…

Boundary DetectionContrastive LearningGeneric Event Boundary Detection