paper-with-me

홈 › Papers

Punching Bag vs. Punching Person: Motion Transferability in Videos

2025-07-31 · Raiyaan Abdullah, Jared Claypoole, Michael Cogswell, Ajay Divakaran, Yogesh Rawat arxiv

Action recognition models demonstrate strong generalization, but can they effectively transfer high-level motion concepts across diverse contexts, even within similar distributions? For example, can a model recognize the broad action "punching" when presented with an unseen variation such as "punching person"? To explore this, we introduce a motion transferability framework with three datasets: (1) Syn-TA, a synthetic dataset with 3D object motions; (2) Kinetics400-TA; and (3) Something-Something-v2-TA, both adapted from natural video datasets. We evaluate 13 state-of-the-art models on these benchmarks and observe a significant drop in performance when recognizing high-level actions in novel contexts. Our analysis reveals: 1) Multimodal models struggle more with fine-grained unknown actions than with coarse ones; 2) The bias-free Syn-TA proves as challenging as real-world datasets, with models showing greater performance drops in controlled settings; 3) Larger models improve transferability when spatial cues dominate but struggle with intensive temporal reasoning, while reliance on object and background cues hinders generalization. We further explore how disentangling coarse and fine motions can improve recognition in temporally challenging datasets. We believe this study establishes a crucial benchmark for assessing motion transferability in action recognition. Datasets and relevant code: https://github.com/raiyaan-abdullah/Motion-Transfer.

📄 PDF Abstract BibTeX arXiv:2508.00085

Code (0)

등록된 구현이 없습니다.

Tasks

Action Recognition

Similar Papers 제목 키워드 기반

First-Person Activity Recognition: What Are They Doing to Me?

2013-06-01 · CVPR 2013 6 · Michael S. Ryoo, Larry Matthies

This paper discusses the problem of recognizing interaction-level human activities from a first-person viewpoint. The goal is to enable an observer (e.g., a robot or a wearable camera) to understand 'what activity others…

Activity RecognitionGeneral Classification

PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions

2026-05-28 · Omer Benishu, Gal Fiebelman, Sagie Benaim arxiv

We address the task of generating physically accurate and visually faithful 4D Human-Object Interaction (HOI). Given a static 3D human and target object represented as 3D Gaussian Splats (3DGS), our goal is to synthesize…

Morphology-Consistent Humanoid Interaction through Robot-Centric Video Synthesis

2026-03-20 · Weisheng Xu, Jian Li, Yi Gu, Bin Yang 외 arxiv

Equipping humanoid robots with versatile interaction skills typically requires either extensive policy training or explicit human-to-robot motion retargeting. However, learning-based policies face prohibitive data collec…

Video GenerationPose Estimation

Reference Grounded Skill Discovery

2025-10-07 · Seungeun Rho, Aaron Trinh, Danfei Xu, Sehoon Ha arxiv

Scaling unsupervised skill discovery algorithms to high-DoF agents remains challenging. As dimensionality increases, the exploration space grows exponentially, while the manifold of meaningful skills remains limited. The…

Real Time Action Recognition from Video Footage

2021-12-13 · Tasnim Sakib Apon, Mushfiqul Islam Chowdhury, MD Zubair Reza, Arpita Datta 외

Crime rate is increasing proportionally with the increasing rate of the population. The most prominent approach was to introduce Closed-Circuit Television (CCTV) camera-based surveillance to tackle the issue. Video surve…

Action Recognition