First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge
In this report, we present our first-place solution to the Multiple-choice Video Question Answering (QA) track of The Second Perception Test Challenge. This competition posed a complex video understanding task, requiring models to accurately comprehend and answer questions about video content. To address this challenge, we leveraged the powerful QwenVL2 (7B) model and fine-tune it on the provided training set. Additionally, we employed model ensemble strategies and Test Time Augmentation to boost performance. Through continuous optimization, our approach achieved a Top-1 Accuracy of 0.7647 on the leaderboard.
Code (0)
등록된 구현이 없습니다.
Tasks
Multiple-choiceQuestion AnsweringVideo Question AnsweringVideo UnderstandingSimilar Papers 제목 키워드 기반
Beyond Pick-and-Place: Tackling Robotic Stacking of Diverse Shapes
We study the problem of robotic stacking with objects of complex geometry. We propose a challenging and diverse set of such objects that was carefully designed to require strategies beyond a simple "pick-and-place" solut…
Offline RLReinforcement Learning (RL)Skill GeneralizationSkill MasteryWinning Amazon KDD Cup'24
This paper describes the winning solution of all 5 tasks for the Amazon KDD Cup 2024 Multi Task Online Shopping Challenge for LLMs. The challenge was to build a useful assistant, answering questions in the domain of onli…
Data AugmentationMultiple-choiceQuantizationSynthetic Data GenerationThe 1st-place Solution for ECCV 2022 Multiple People Tracking in Group Dance Challenge
We present our 1st place solution to the Group Dance Multiple People Tracking Challenge. Based on MOTR: End-to-End Multiple-Object Tracking with Transformer, we explore: 1) detect queries as anchors, 2) tracking as query…
Multi-Object TrackingMultiple Object TrackingMultiple Object Tracking with TransformerMultiple People TrackingA Nuclear-norm Model for Multi-Frame Super-Resolution Reconstruction from Video Clips
We propose a variational approach to obtain super-resolution images from multiple low-resolution frames extracted from video clips. First the displacement between the low-resolution frames and the reference frame are com…
Multi-Frame Super-ResolutionOptical Flow EstimationSuper-ResolutionThe 1st Solution for 4th PVUW MeViS Challenge: Unleashing the Potential of Large Multimodal Models for Referring Video Segmentation
Motion expression video segmentation is designed to segment objects in accordance with the input motion expressions. In contrast to the conventional Referring Video Object Segmentation (RVOS), it places emphasis on motio…
Inference OptimizationReferring Video Object SegmentationSegmentationSemantic Segmentation+3