paper-with-me

Papers Action Classification

“Action Classification” 태그가 달린 논문 485편 · 필터 해제

P-JEPA: Procedural Video Representation Learning via Joint Embedding Predictive Architecture

2026-06-22 · Felix Tristram, Stefano Gasperini, Benjamin Killeen, Marcel Walch 외 arxiv

The increasing maturity of embodied AI platforms has driven a growing interest in procedural video representation learning to support intelligent assistance systems for complex, multi-step tasks. Leveraging large-scale l…

Representation LearningAction ClassificationAction Segmentation

T-MOR: Learning Motion-Aware Skeleton Representations for Human Action Recognition

2026-06-19 · Di Yang, Mahmoud Ali, Quan Kong, Gianpiero Francesca 외 arxiv

Vision-language models such as CLIP have recently achieved strong performance on a wide range of visual understanding tasks. However, most existing models rely primarily on appearance-level supervision from images or vid…

Action ClassificationContrastive LearningAction UnderstandingAction Recognition

Decoupled Object-Centric Video Understanding for Generating Robotic Manipulation Commands

2026-06-15 · Thanh Nguyen Canh, Thanh-Tuan Tran, Haolan Zhang, Ziyan Gao 외 arxiv

Translating video demonstrations into executable robot commands remains challenging because existing methods often fail to identify which objects are functionally involved in the demonstrated action. As a result, they ma…

Zero-shot GeneralizationAction ClassificationAction Recognition

PlayClass: Automated Play Behaviour Classification in Poultry

2026-05-26 · Prince Ravi Leow, Neil Scheidwasser, Rebecca Oscarsson, Per Jensen 외 arxiv

Automated monitoring of animal welfare has largely targeted negative indicators, leaving positive welfare behaviours such as play underexplored. To address this gap, we present PlayClass, a pipeline for play-behaviour cl…

Action Classification

Zero-Shot Temporal Action Localization Through Textual Guidance

2026-05-21 · Benedetta Liberatori, Alessandro Conti, Lorenzo Vaquero, Paolo Rota 외 arxiv

Zero-shot temporal action localization (ZS-TAL) consists of classifying and localizing actions in untrimmed videos, where action classes are unseen at training time. Existing work uses Vision and Language Models (VLMs), …

Temporal Action LocalizationAction Classification

ATRACT: A Trustworthy Robotic Autonomous system to support Casualty Triage

2026-05-16 · Tasweer Ahmad, Rafael Pina, Sandip Pradhan, Arindam Sikdar 외 arxiv

At a time when drones are increasingly associated with hostile operations, we re-purpose them for humanitarian and life-saving applications. However, adapting search and rescue drones for battlefield triage remains extre…

Action ClassificationData Augmentation

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders

2026-04-05 · Atahan Dokme, Sriram Vishwanath arxiv

We present the first systematic study of Sparse Autoencoders (SAEs) on video representations. Standard SAEs decompose video into interpretable, monosemantic features but destroy temporal coherence: hard TopK selection pr…

Action ClassificationVideo Retrieval

ViTs for Action Classification in Videos: An Approach to Risky Tackle Detection in American Football Practice Videos

2026-04-01 · Syed Ahsan Masud Zaidi, William Hsu, Scott Dietrich arxiv

Early identification of hazardous actions in contact sports enables timely intervention and improves player safety. We present a method for detecting risky tackles in American football practice videos and introduce a sub…

Action Classification

Shared Representation for 3D Pose Estimation, Action Classification, and Progress Prediction from Tactile Signals

2026-03-26 · Isaac Han, Seoyoung Lee, Sangyeon Park, Ecehan Akan 외 arxiv

Estimating human pose, classifying actions, and predicting movement progress are essential for human-robot interaction. While vision-based methods suffer from occlusion and privacy concerns in realistic environments, tac…

3D Human Pose EstimationAction ClassificationMulti-Task Learning3D Pose Estimation

WiFi2Cap: Semantic Action Captioning from Wi-Fi CSI via Limb-Level Semantic Alignment

2026-03-24 · Tzu-Ti Wei, Chu-Yu Huang, Yu-Chee Tseng, Jen-Jee Chen arxiv

Privacy-preserving semantic understanding of human activities is important for indoor sensing, yet existing Wi-Fi CSI-based systems mainly focus on pose estimation or predefined action classification rather than fine-gra…

Action ClassificationPose Estimation

Human-to-Robot Interaction: Learning from Video Demonstration for Robot Imitation

2026-02-22 · Thanh Nguyen Canh, Thanh-Tuan Tran, Haolan Zhang, Ziyan Gao 외 arxiv

Learning from Demonstration (LfD) offers a promising paradigm for robot skill acquisition. Recent approaches attempt to extract manipulation commands directly from video demonstrations, yet face two critical challenges: …

Reinforcement LearningAction ClassificationRobot ManipulationVideo Captioning

Autonomous Action Runtime Management(AARM):A System Specification for Securing AI-Driven Actions at Runtime

2026-02-10 · Herman Errico arxiv

As artificial intelligence systems evolve from passive assistants into autonomous agents capable of executing consequential actions, the security boundary shifts from model outputs to tool execution. Traditional security…

Action Classification

Two-Stream temporal transformer for video action classification

2026-01-20 · Nattapong Kurpukdee, Adrian G. Bors arxiv

Motion representation plays an important role in video understanding and has many applications including action recognition, robot and autonomous guidance or others. Lately, transformer networks, through their self-atten…

Action ClassificationAction Recognition

Explainable Action Form Assessment by Exploiting Multimodal Chain-of-Thoughts Reasoning

2025-12-17 · Mengshi Qi, Yeteng Wu, Wulian Yun, Xianlin Zhang 외 arxiv

Evaluating whether human action is standard or not and providing reasonable feedback to improve action standardization is very crucial but challenging in real-world scenarios. However, current video understanding methods…

Action Quality AssessmentExplanation GenerationAction Classification

Recurrent Video Masked Autoencoders

2025-12-15 · Daniel Zoran, Nikhil Parthasarathy, Yi Yang, Drew A Hudson 외 arxiv

We present Recurrent Video Masked-Autoencoders (RVM): a novel approach to video representation learning that leverages recurrent computation to model the temporal structure of video data. RVM couples an asymmetric maskin…

Representation LearningKnowledge DistillationAction ClassificationObject Tracking

Improving action classification with brain-inspired deep networks

2025-12-08 · Aidas Aglinskas, Stefano Anzellotti arxiv

Action recognition is also key for applications ranging from robotics to healthcare monitoring. Action information can be extracted from the body pose and movements, as well as from the background scene. However, the ext…

Action ClassificationAction Recognition

DisMo: Disentangled Motion Representations for Open-World Motion Transfer

2025-11-28 · Thomas Ressler-Antal, Frank Fundel, Malek Ben Alaya, Stefan Andreas Baumann 외 arxiv

Recent advances in text-to-video (T2V) and image-to-video (I2V) models, have enabled the creation of visually compelling and dynamic videos from simple textual descriptions or initial frames. However, these models often …

Action Classification

Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations

2025-11-25 · Chao Wang, Chengan Che, Xinyue Chen, Sophia Tsoka 외 arxiv

Counterfactual explanations (CFEs) are minimal and semantically meaningful modifications of the input of a model that alter the model predictions. They highlight the decisive features the model relies on, providing contr…

Emotion ClassificationAction Classification

Predicting Social Media User Actions: A Hybrid Approach for Common and Rare Behavior Prediction on Bluesky

2025-11-21 · Benjamin White, Anastasia Shimorina arxiv

Understanding and predicting user behavior on social media platforms is crucial for content recommendation and platform design. While existing approaches focus primarily on common actions like retweeting and liking, the …

Action Classification

SUGAR: Learning Skeleton Representation with Visual-Motion Knowledge for Action Recognition

2025-11-13 · Qilang Ye, Yu Zhou, Lian He, Jie Zhang 외 arxiv

Large Language Models (LLMs) hold rich implicit knowledge and powerful transferability. In this paper, we explore the combination of LLMs with the human skeleton to perform action classification and description. However,…

Action ClassificationAction Recognition
1–20 / 485 다음 →