Papers Video Classification
“Video Classification” 태그가 달린 논문 495편 · 필터 해제
Learning to Unify Deformable Shape and Texture Representations for Cardiac Video Classification
Deformable shape representations have proven to be robust complements to texture features in cardiac image classification, offering geometric priors that are invariant to imaging artifacts and intensity variations. Howev…
Image ClassificationVideo ClassificationGen4U: Unifying Video Generation and Understanding via Diffusion
Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models overcome this limitation. By systematicall…
Camera Pose EstimationVideo ClassificationDepth EstimationVideo CaptioningForget, Anticipate and Adapt: Test Time Training for Long Videos
Test Time Training (TTT) is a mechanism in which a model adapts to an incoming test-sample by performing some self-supervised (SSL) task and updating its weights even during inference. This procedure does not require lab…
Video ClassificationSpatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos
Automated classification of standard echocardiographic views is crucial for efficient clinical workflow but faces three main challenges. First, publicly available datasets are scarce and limited in scale and view coverag…
Video ClassificationInterpretable Temporal Facial-Region Motion Analysis for In-the-Wild Parkinson's Disease Video Classification
Reduced facial expressivity is a common motor manifestation of Parkinson's disease (PD), often described as hypomimia or facial bradykinesia. This paper examines whether temporal motion descriptors extracted from facial-…
Binary ClassificationVideo ClassificationArtifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos
Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inconsistencies, structural distortions, and semantic incoherence. While…
Video ClassificationArtifact DetectionAppearance-free Action Recognition: Zero-shot Generalization in Humans and a Two-Pathway Model
Action recognition is a fundamental ability for social species. Yet, its underlying computations are not well understood. Classical psychophysical studies using simplified stimuli have shown that humans can perceive body…
Zero-shot GeneralizationVideo ClassificationAction RecognitionLOLGORITHM: Funny Comment Generation Agent For Short Videos
Short-form video platforms have become central to multimedia information dissemination, where comments play a critical role in driving engagement, propagation, and algorithmic feedback. However, existing approaches -- in…
Video ClassificationVideo SummarizationSemantic RetrievalvAccSOL: Efficient and Transparent AI Vision Offloading for Mobile Robots
Mobile robots are increasingly deployed for inspection, patrol, and search-and-rescue operations, relying on computer vision for perception, navigation, and autonomous decision-making. However, executing modern vision wo…
Semantic SegmentationVideo ClassificationImage ClassificationHSEmotion Team at ABAW-10 Competition: Facial Expression Recognition, Valence-Arousal Estimation, Action Unit Detection and Fine-Grained Violence Classification
This article presents our results for the 10th Affective Behavior Analysis in-the-Wild (ABAW) competition. For frame-wise facial emotion understanding tasks (frame-wise facial expression recognition, valence-arousal esti…
Facial Expression RecognitionAction Unit DetectionVideo ClassificationEmotion RecognitionFrom Imitation to Intuition: Intrinsic Reasoning for Open-Instance Video Classification
Conventional video classification models, acting as effective imitators, excel in scenarios with homogeneous data distributions. However, real-world applications often present an open-instance challenge, where intra-clas…
Reinforcement LearningVideo ClassificationContrastive learning-based video quality assessment-jointed video vision transformer for video recognition
Video quality significantly affects video classification. We found this problem when we classified Mild Cognitive Impairment well from clear videos, but worse from blurred ones. From then, we realized that referring to V…
Video Quality AssessmentSelf-Supervised LearningVideo ClassificationContrastive LearningAssessing Situational and Spatial Awareness of VLMs with Synthetically Generated Video
Spatial reasoning in vision language models (VLMs) remains fragile when semantics hinge on subtle temporal or geometric cues. We introduce a synthetic benchmark that probes two complementary skills: situational awareness…
Video ClassificationSpatial ReasoningNot all Blends are Equal: The BLEMORE Dataset of Blended Emotion Expressions with Relative Salience Annotations
Humans often experience not just a single basic emotion at a time, but rather a blend of several emotions with varying salience. Despite the importance of such blended emotions, most video-based emotion recognition appro…
Video ClassificationEmotion RecognitionCell Behavior Video Classification Challenge, a benchmark for computer vision methods in time-lapse microscopy
The classification of microscopy videos capturing complex cellular behaviors is crucial for understanding and quantifying the dynamics of biological processes over time. However, it remains a frontier in computer vision,…
Video ClassificationEffects of Different Attention Mechanisms Applied on 3D Models in Video Classification
Human action recognition has become an important research focus in computer vision due to the wide range of applications where it is used. 3D Resnet-based CNN models, particularly MC3, R3D, and R(2+1)D, have different co…
Video ClassificationAction RecognitionVL-JEPA: Joint Embedding Predictive Architecture for Vision-language
We introduce VL-JEPA, a vision-language model built on a Joint Embedding Predictive Architecture (JEPA). Instead of autoregressively generating tokens as in classical VLMs, VL-JEPA predicts continuous embeddings of the t…
Video ClassificationVideo RetrievalOmniFD: A Unified Model for Versatile Face Forgery Detection
Face forgery detection encompasses multiple critical tasks, including identifying forged images and videos and localizing manipulated regions and temporal segments. Current approaches typically employ task-specific model…
Video ClassificationMulti-Task LearningRB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification
Vision Language Models (VLMs) are becoming increasingly integral to multimedia understanding; however, they often struggle with domain-specific video classification tasks, particularly in cases with limited data. This st…
Video ClassificationAuto-US: An Ultrasound Video Diagnosis Agent Using Video Classification Framework and LLMs
AI-assisted ultrasound video diagnosis presents new opportunities to enhance the efficiency and accuracy of medical imaging analysis. However, existing research remains limited in terms of dataset diversity, diagnostic p…
Video Classification