paper-with-me

VidSTG

홈페이지 · 논문 29편

The VidSTG dataset is a spatio-temporal video grounding dataset constructed based on the video relation dataset VidOR. VidOR contains 7,000, 835 and 2,165 videos for training, validation and testing, respectively. The goal of the Spatio-Temporal Video Grounding task (STVG) is to localize the spatio-temporal section of an untrimmed video that matches a given sentence depicting an object. VidSTG contains 5,563, 618, and 743 videos for training, validation, and testing, respectively. Source: https://github.com/Guaranteer/VidSTG-Dataset Image Source: https://github.com/Guaranteer/VidSTG-Dataset

벤치마크

Spatio-Temporal Video Grounding on VidSTG 결과 3개