Video Classification
12개 벤치마크 · 논문 495편 · 이 태스크의 논문 보기 →
Benchmarks
Breakfast
COIN
MoB
YouTube-8M
Charades
Home Action Genome
Kinetics
Multimodal PISA
Something-Something V1
Something-Something V2
Most implemented
Non-local Neural Networks
Group Normalization
Is Space-Time Attention All You Need for Video Understanding?
Video Swin Transformer
Temporal Segment Networks for Action Recognition in Videos
Papers
Learning to Unify Deformable Shape and Texture Representations for Cardiac Video Classification
Deformable shape representations have proven to be robust complements to texture features in cardiac image classification, offering geometric priors that are invariant to imaging artifacts and intensity variations. Howev…
Image ClassificationVideo ClassificationGen4U: Unifying Video Generation and Understanding via Diffusion
Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models overcome this limitation. By systematicall…
Camera Pose EstimationVideo ClassificationDepth EstimationVideo CaptioningForget, Anticipate and Adapt: Test Time Training for Long Videos
Test Time Training (TTT) is a mechanism in which a model adapts to an incoming test-sample by performing some self-supervised (SSL) task and updating its weights even during inference. This procedure does not require lab…
Video ClassificationSpatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos
Automated classification of standard echocardiographic views is crucial for efficient clinical workflow but faces three main challenges. First, publicly available datasets are scarce and limited in scale and view coverag…
Video ClassificationInterpretable Temporal Facial-Region Motion Analysis for In-the-Wild Parkinson's Disease Video Classification
Reduced facial expressivity is a common motor manifestation of Parkinson's disease (PD), often described as hypomimia or facial bradykinesia. This paper examines whether temporal motion descriptors extracted from facial-…
Binary ClassificationVideo ClassificationArtifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inconsistencies, structural distortions, and semantic incoherence. While…
Video ClassificationArtifact Detection