Technical Report for Soccernet 2023 -- Dense Video Captioning
In the task of dense video captioning of Soccernet dataset, we propose to generate a video caption of each soccer action and locate the timestamp of the caption. Firstly, we apply Blip as our video caption framework to generate video captions. Then we locate the timestamp by using (1) multi-size sliding windows (2) temporal proposal generation and (3) proposal classification.
Code (0)
등록된 구현이 없습니다.
Tasks
Dense Video CaptioningVideo CaptioningMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
SoccerNet 2024 Challenges Results
The SoccerNet 2024 challenges represent the fourth annual video understanding challenges organized by the SoccerNet team. These challenges aim to advance research across multiple themes in football, including broadcast v…
Action SpottingDense Video CaptioningGame State ReconstructionVideo Captioning+1Dense-Captioning Events in Videos: SYSU Submission to ActivityNet Challenge 2020
This technical report presents a brief description of our submission to the dense video captioning task of ActivityNet Challenge 2020. Our approach follows a two-stage pipeline: first, we extract a set of temporal event …
Dense CaptioningDense Video CaptioningVideo CaptioningTeam RUC_AIM3 Technical Report at Activitynet 2020 Task 2: Exploring Sequential Events Detection for Dense Video Captioning
Detecting meaningful events in an untrimmed video is essential for dense video captioning. In this work, we propose a novel and simple model for event sequence generation and explore temporal relationships of the event s…
Dense CaptioningDense Video CaptioningTask 2Video CaptioningAction Spotting using Dense Detection Anchors Revisited: Submission to the SoccerNet Challenge 2022
This brief technical report describes our submission to the Action Spotting SoccerNet Challenge 2022. The challenge was part of the CVPR 2022 ActivityNet Workshop. Our submission was based on a recently proposed method w…
Action SpottingSoccerNet 2023 Challenges Results
The SoccerNet 2023 challenges were the third annual video understanding challenges organized by the SoccerNet team. For this third edition, the challenges were composed of seven vision-based tasks split into three main t…
Action SpottingCamera CalibrationDense Video CaptioningJersey Number Recognition+4