Action Recognition
56개 벤치마크 · 논문 3,044편 · 이 태스크의 논문 보기 →
Benchmarks
Something-Something V2
UCF101
HMDB-51
Something-Something V1
AVA v2.2
EPIC-KITCHENS-100
NTU RGB+D
NTU RGB+D 120
Diving-48
ActivityNet
AVA v2.1
H2O (2 Hands and Objects)
THUMOS’14
Sports-1M
HACS
Charades-Ego
Animal Kingdom
BAR
HAA500
LoTE-Animal
UAV-Human
Volleyball
RareAct
UCF-101
Drone-Action
ICVL-4
IRD
Mimetics
Okutama-Action
Penn Action
SL-Animals
miniSports
HMDB51
ActionNet-VE
Charades
DVS128 Gesture
EPIC-KITCHENS-55
EgoGesture
Hockey
IndustReal
KTH
MECCANO
MTL-AQA
N-UCLA
NEC Drone
RoCoG-v2
Skeleton-Mimetics
THUMOS14
UAV Human
UCF 101
UCFSports
UTD-MHAD
VIRAT Ground 2.0
Most implemented
EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks
Learning Transferable Visual Models From Natural Language Supervision
Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset
Non-local Neural Networks
Learning Spatiotemporal Features with 3D Convolutional Networks
Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet?
Papers
SAFER-Activities: A Dataset for Smart Assessment of Fall Events and Routine Activities
Smart healthcare monitoring systems require precise action recognition to ensure well-being and timely intervention in critical situations such as falls, particularly for mobility-challenged individuals. Existing dataset…
Action RecognitionFew-Shot Video Recognition via Hierarchical Metric Learning
Few-shot action recognition (FSAR) aims to recognize unseen action categories with only a small number of annotated video samples. Recent works typically apply single-prototype supervision at the network output and fail …
Action RecognitionMetric LearningHidden In Plain Gaze: Gaze Representations as Privacy Controls for Utility and Re-identification Risk in XR
Intelligent extended reality (XR) systems increasingly use eye and head tracking to infer user intent, task, and attention, but the same signals can also reveal biometric identity. We study whether gaze data representati…
Action RecognitionPose-Anchored Optical Flow for Low-Latency Human Action Anticipation in Human-Robot Teaming
Human-robot interaction (HRI) requires robots to interpret human actions early in their execution in order to respond safely, efficiently, and naturally. However, many existing approaches to human action recognition rely…
Action AnticipationAction RecognitionMoTE: Mixture of Task Experts for Multi-Task Video Understanding
Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recognition, forecasting, and procedure prediction. Dense transformer decoders share the same feed-forward n…
Action RecognitionByteAction: Byte-space Action Recognition Foundation Model
Byte-space Action Recognition (BAR) aims to recognize human actions directly from compressed image bitstreams without any pixel decoding. By operating entirely in byte space, BAR is inherently independent of file integri…
Action Recognition