Technical Report for CVPR 2022 LOVEU AQTC Challenge
This technical report presents the 2nd winning model for AQTC, a task newly introduced in CVPR 2022 LOng-form VidEo Understanding (LOVEU) challenges. This challenge faces difficulties with multi-step answers, multi-modal, and diverse and changing button representations in video. We address this problem by proposing a new context ground module attention mechanism for more effective feature mapping. In addition, we also perform the analysis over the number of buttons and ablation study of different step networks and video features. As a result, we achieved the overall 2nd place in LOVEU competition track 3, specifically the 1st place in two out of four evaluation metrics. Our code is available at https://github.com/jaykim9870/ CVPR-22_LOVEU_unipyler.
Code (1)
Tasks
Video UnderstandingSimilar Papers 제목 키워드 기반
Winning the CVPR'2022 AQTC Challenge: A Two-stage Function-centric Approach
Affordance-centric Question-driven Task Completion for Egocentric Assistant(AQTC) is a novel task which helps AI assistant learn from instructional videos and scripts and guide the user step-by-step. In this paper, we de…
A Solution to CVPR'2023 AQTC Challenge: Video Alignment for Multi-Step Inference
Affordance-centric Question-driven Task Completion (AQTC) for Egocentric Assistant introduces a groundbreaking scenario. In this scenario, through learning instructional videos, AI assistants provide users with step-by-s…
Video AlignmentFirst Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment
Affordance-Centric Question-driven Task Completion (AQTC) has been proposed to acquire knowledge from videos to furnish users with comprehensive and systematic instructions. However, existing methods have hitherto neglec…
Human-Object Interaction DetectionTechnical Report for CVPR 2024 WeatherProof Dataset Challenge: Semantic Segmentation on Paired Real Data
This technical report presents the implementation details of 2nd winning for CVPR'24 UG2 WeatherProof Dataset Challenge. This challenge aims at semantic segmentation of images degraded by various degrees of weather from …
Semantic SegmentationWinning the CVPR'2021 Kinetics-GEBD Challenge: Contrastive Learning Approach
Generic Event Boundary Detection (GEBD) is a newly introduced task that aims to detect "general" event boundaries that correspond to natural human perception. In this paper, we introduce a novel contrastive learning base…
Boundary DetectionContrastive LearningGeneric Event Boundary Detection