paper-with-me

Papers Video Classification

“Video Classification” 태그가 달린 논문 495편 · 필터 해제

Learning to Unify Deformable Shape and Texture Representations for Cardiac Video Classification

2026-07-08 · Tonmoy Hossain, Miaomiao Zhang arxiv

Deformable shape representations have proven to be robust complements to texture features in cardiac image classification, offering geometric priors that are invariant to imaging artifacts and intensity variations. Howev…

Image ClassificationVideo Classification

Gen4U: Unifying Video Generation and Understanding via Diffusion

2026-07-07 · Michael King, Aravindh Mahendran, Matthew Koichi Grimes, Fedor Kitashov 외 arxiv

Prior work suggests that diffusion representations capture low-level geometry but struggle with high-level semantics. We demonstrate that state-of-the-art video diffusion models overcome this limitation. By systematicall…

Camera Pose EstimationVideo ClassificationDepth EstimationVideo Captioning

Forget, Anticipate and Adapt: Test Time Training for Long Videos

2026-06-25 · Rajat Modi, Sebastian Noel, Xin Liang, Yogesh Singh Rawat arxiv

Test Time Training (TTT) is a mechanism in which a model adapts to an incoming test-sample by performing some self-supervised (SSL) task and updating its weights even during inference. This procedure does not require lab…

Video Classification

Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos

2026-06-16 · Bo Gou, Jicheng Zhang, Jianlong Xiong, Tao He 외 arxiv

Automated classification of standard echocardiographic views is crucial for efficient clinical workflow but faces three main challenges. First, publicly available datasets are scarce and limited in scale and view coverag…

Video Classification

Interpretable Temporal Facial-Region Motion Analysis for In-the-Wild Parkinson's Disease Video Classification

2026-06-08 · Riyadh Almushrafy arxiv

Reduced facial expressivity is a common motor manifestation of Parkinson's disease (PD), often described as hypomimia or facial bradykinesia. This paper examines whether temporal motion descriptors extracted from facial-…

Binary ClassificationVideo Classification

Artifact-Bench: Evaluating MLLMs on Detecting and Assessing the Artifacts of AI-Generated Videos

2026-05-18 · Yuqi Tang, Yang Shi, Zhuoran Zhang, Qixun Wang 외 arxiv

Recent video generative models have greatly improved the realism of AI-generated videos, yet their outputs still exhibit artifacts such as temporal inconsistencies, structural distortions, and semantic incoherence. While…

Video ClassificationArtifact Detection

Appearance-free Action Recognition: Zero-shot Generalization in Humans and a Two-Pathway Model

2026-04-17 · Prerana Kumar, Martin A. Giese arxiv

Action recognition is a fundamental ability for social species. Yet, its underlying computations are not well understood. Classical psychophysical studies using simplified stimuli have shown that humans can perceive body…

Zero-shot GeneralizationVideo ClassificationAction Recognition

LOLGORITHM: Funny Comment Generation Agent For Short Videos

2026-04-09 · Xuan Ouyang, Bouzhou Wang, Senan Wang, Siyuan Xiahou 외 arxiv

Short-form video platforms have become central to multimedia information dissemination, where comments play a critical role in driving engagement, propagation, and algorithmic feedback. However, existing approaches -- in…

Video ClassificationVideo SummarizationSemantic Retrieval

vAccSOL: Efficient and Transparent AI Vision Offloading for Mobile Robots

2026-03-17 · Adam Zahir, Michele Gucciardom Falk Selker, Anastasios Nanos, Kostis Papazafeiropoulos 외 arxiv

Mobile robots are increasingly deployed for inspection, patrol, and search-and-rescue operations, relying on computer vision for perception, navigation, and autonomous decision-making. However, executing modern vision wo…

Semantic SegmentationVideo ClassificationImage Classification

HSEmotion Team at ABAW-10 Competition: Facial Expression Recognition, Valence-Arousal Estimation, Action Unit Detection and Fine-Grained Violence Classification

2026-03-13 · Andrey V. Savchenko, Kseniia Tsypliakova arxiv

This article presents our results for the 10th Affective Behavior Analysis in-the-Wild (ABAW) competition. For frame-wise facial emotion understanding tasks (frame-wise facial expression recognition, valence-arousal esti…

Facial Expression RecognitionAction Unit DetectionVideo ClassificationEmotion Recognition

From Imitation to Intuition: Intrinsic Reasoning for Open-Instance Video Classification

2026-03-11 · Ke Zhang, Xiangchen Zhao, Yunjie Tian, Jiayu Zheng 외 arxiv

Conventional video classification models, acting as effective imitators, excel in scenarios with homogeneous data distributions. However, real-world applications often present an open-instance challenge, where intra-clas…

Reinforcement LearningVideo Classification

Contrastive learning-based video quality assessment-jointed video vision transformer for video recognition

2026-03-11 · Jian Sun, Mohammad H. Mahoor arxiv

Video quality significantly affects video classification. We found this problem when we classified Mild Cognitive Impairment well from clear videos, but worse from blurred ones. From then, we realized that referring to V…

Video Quality AssessmentSelf-Supervised LearningVideo ClassificationContrastive Learning

Assessing Situational and Spatial Awareness of VLMs with Synthetically Generated Video

2026-01-22 · Pascal Benschop, Justin Dauwels, Jan van Gemert arxiv

Spatial reasoning in vision language models (VLMs) remains fragile when semantics hinge on subtle temporal or geometric cues. We introduce a synthetic benchmark that probes two complementary skills: situational awareness…

Video ClassificationSpatial Reasoning

Not all Blends are Equal: The BLEMORE Dataset of Blended Emotion Expressions with Relative Salience Annotations

2026-01-19 · Tim Lachmann, Alexandra Israelsson, Christina Tornberg, Teimuraz Saghinadze 외 arxiv

Humans often experience not just a single basic emotion at a time, but rather a blend of several emotions with varying salience. Despite the importance of such blended emotions, most video-based emotion recognition appro…

Video ClassificationEmotion Recognition

Cell Behavior Video Classification Challenge, a benchmark for computer vision methods in time-lapse microscopy

2026-01-15 · Raffaella Fiamma Cabini, Deborah Barkauskas, Guangyu Chen, Zhi-Qi Cheng 외 arxiv

The classification of microscopy videos capturing complex cellular behaviors is crucial for understanding and quantifying the dynamics of biological processes over time. However, it remains a frontier in computer vision,…

Video Classification

Effects of Different Attention Mechanisms Applied on 3D Models in Video Classification

2026-01-15 · Mohammad Rasras, Iuliana Marin, Serban Radu, Irina Mocanu arxiv

Human action recognition has become an important research focus in computer vision due to the wide range of applications where it is used. 3D Resnet-based CNN models, particularly MC3, R3D, and R(2+1)D, have different co…

Video ClassificationAction Recognition

VL-JEPA: Joint Embedding Predictive Architecture for Vision-language

2025-12-11 · Delong Chen, Mustafa Shukor, Theo Moutakanni, Willy Chung 외 arxiv

We introduce VL-JEPA, a vision-language model built on a Joint Embedding Predictive Architecture (JEPA). Instead of autoregressively generating tokens as in classical VLMs, VL-JEPA predicts continuous embeddings of the t…

Video ClassificationVideo Retrieval

OmniFD: A Unified Model for Versatile Face Forgery Detection

2025-11-30 · Haotian Liu, Haoyu Chen, Chenhui Pan, You Hu 외 arxiv

Face forgery detection encompasses multiple critical tasks, including identifying forged images and videos and localizing manipulated regions and temporal segments. Current approaches typically employ task-specific model…

Video ClassificationMulti-Task Learning

RB-FT: Rationale-Bootstrapped Fine-Tuning for Video Classification

2025-11-19 · Meilong Xu, Di Fu, Jiaxing Zhang, Gong Yu 외 arxiv

Vision Language Models (VLMs) are becoming increasingly integral to multimedia understanding; however, they often struggle with domain-specific video classification tasks, particularly in cases with limited data. This st…

Video Classification

Auto-US: An Ultrasound Video Diagnosis Agent Using Video Classification Framework and LLMs

2025-11-11 · Yuezhe Yang, Yiyue Guo, Wenjie Cai, Qingqing Ruan 외 arxiv

AI-assisted ultrasound video diagnosis presents new opportunities to enhance the efficiency and accuracy of medical imaging analysis. However, existing research remains limited in terms of dataset diversity, diagnostic p…

Video Classification
1–20 / 495 다음 →