paper-with-me

Video Classification

12개 벤치마크 · 논문 495편 · 이 태스크의 논문 보기 →

Benchmarks

Breakfast

결과 18개

COIN

결과 14개

MoB

결과 6개

YouTube-8M

결과 6개

Charades

결과 2개

Home Action Genome

결과 2개

Kinetics

결과 2개

Multimodal PISA

결과 2개

Most implemented

Non-local Neural Networks

2017-11-21 · 구현 32개

Group Normalization

2018-03-22 · 구현 22개

Video Swin Transformer

2021-06-24 · 구현 15개

Papers

Learning to Unify Deformable Shape and Texture Representations for Cardiac Video Classification

2026-07-08 · Tonmoy Hossain, Miaomiao Zhang arxiv

Deformable shape representations have proven to be robust complements to texture features in cardiac image classification, offering geometric priors that are invariant to imaging artifacts and intensity variations. Howev…

Image ClassificationVideo Classification

Gen4U: Unifying Video Generation and Understanding via Diffusion

2026-07-07 · Michael King, Aravindh Mahendran, Matthew Koichi Grimes, Fedor Kitashov 외 arxiv

Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models overcome this limitation. By systematicall…

Camera Pose EstimationVideo ClassificationDepth EstimationVideo Captioning

Forget, Anticipate and Adapt: Test Time Training for Long Videos

2026-06-25 · Rajat Modi, Sebastian Noel, Xin Liang, Yogesh Singh Rawat arxiv

Test Time Training (TTT) is a mechanism in which a model adapts to an incoming test-sample by performing some self-supervised (SSL) task and updating its weights even during inference. This procedure does not require lab…

Video Classification

Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos

2026-06-16 · Bo Gou, Jicheng Zhang, Jianlong Xiong, Tao He 외 arxiv

Automated classification of standard echocardiographic views is crucial for efficient clinical workflow but faces three main challenges. First, publicly available datasets are scarce and limited in scale and view coverag…

Video Classification

Interpretable Temporal Facial-Region Motion Analysis for In-the-Wild Parkinson's Disease Video Classification

2026-06-08 · Riyadh Almushrafy arxiv

Reduced facial expressivity is a common motor manifestation of Parkinson's disease (PD), often described as hypomimia or facial bradykinesia. This paper examines whether temporal motion descriptors extracted from facial-…

Binary ClassificationVideo Classification

Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

2026-05-18 · Yuqi Tang, Yang Shi, Zhuoran Zhang, Qixun Wang 외 arxiv

Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inconsistencies, structural distortions, and semantic incoherence. While…

Video ClassificationArtifact Detection

전체 495편 보기 →