paper-with-me

ActivityNet

홈페이지 · 논문 807편

The ActivityNet dataset contains 200 different types of activities and a total of 849 hours of videos collected from YouTube. ActivityNet is the largest benchmark for temporal activity detection to date in terms of both the number of activity categories and number of videos, making the task particularly challenging. Version 1.3 of the dataset contains 19994 untrimmed videos in total and is divided into three disjoint subsets, training, validation, and testing by a ratio of 2:1:1. On average, each activity category has 137 untrimmed videos. Each video on average has 1.41 activities which are annotated with temporal boundaries. The ground-truth annotations of test videos are not public. Source: Dynamic Temporal Pyramid Network: A Closer Look at Multi-Scale Modeling for Activity Detection

Videos

벤치마크

Temporal Action Localization on ActivityNet-1.3 결과 132개
Weakly Supervised Action Localization on ActivityNet-1.2 결과 95개
Weakly Supervised Action Localization on ActivityNet-1.3 결과 85개
Video Retrieval on ActivityNet 결과 62개
Temporal Action Proposal Generation on ActivityNet-1.3 결과 44개
Action Recognition on ActivityNet 결과 33개
Zero-Shot Video Retrieval on ActivityNet 결과 12개
GZSL Video Classification on ActivityNet-GZSL(main) 결과 7개
Weakly-supervised Temporal Action Localization on ActivityNet-1.3 결과 5개
Zero-Shot Action Recognition on ActivityNet 결과 5개
GZSL Video Classification on ActivityNet-GZSL (cls) 결과 4개
Temporal Action Localization on ActivityNet-1.2 결과 4개
Action Classification on ActivityNet-1.2 결과 3개
Action Recognition In Videos on ActivityNet 결과 3개
Action Classification on ActivityNet 결과 1개
Few Shot Temporal Action Localization on ActivityNet 결과 1개
Visual Question Answering (VQA) on ActivityNet 결과 1개