paper-with-me

홈 › Papers

SkillFormer: Unified Multi-View Video Understanding for Proficiency Estimation

2025-05-13 · Edoardo Bianchi, Antonio Liotta

Assessing human skill levels in complex activities is a challenging problem with applications in sports, rehabilitation, and training. In this work, we present SkillFormer, a parameter-efficient architecture for unified multi-view proficiency estimation from egocentric and exocentric videos. Building on the TimeSformer backbone, SkillFormer introduces a CrossViewFusion module that fuses view-specific features using multi-head cross-attention, learnable gating, and adaptive self-calibration. We leverage Low-Rank Adaptation to fine-tune only a small subset of parameters, significantly reducing training costs. In fact, when evaluated on the EgoExo4D dataset, SkillFormer achieves state-of-the-art accuracy in multi-view settings while demonstrating remarkable computational efficiency, using 4.5x fewer parameters and requiring 3.75x fewer training epochs than prior baselines. It excels in multiple structured tasks, confirming the value of multi-view integration for fine-grained skill assessment.

📄 PDF Abstract BibTeX arXiv:2505.08665

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyVideo Understanding

Methods 이 논문이 사용한 방법론

TimeSformer TimeSformer is a convolution-free approach to video classification built exclusively on self-attention over space and time. It…

Similar Papers 제목 키워드 기반

PATS: Proficiency-Aware Temporal Sampling for Multi-View Sports Skill Assessment

2025-06-05 · Edoardo Bianchi, Antonio Liotta

Automated sports skill assessment requires capturing fundamental movement patterns that distinguish expert from novice performance, yet current video sampling methods disrupt the temporal continuity essential for profici…

Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback

2026-05-05 · Edoardo Bianchi, Antonio Liotta arxiv

Estimating how well a person performs an action, rather than which action is performed, is central to coaching, rehabilitation, and talent identification. This task is challenging because proficiency is encoded in subtle…

Video Understanding: From Geometry and Semantics to Unified Models

2026-03-18 · Zhaochong An, Zirui Li, Mingqiao Ye, Feng Qiao 외 arxiv

Video understanding aims to enable models to perceive, reason about, and interact with the dynamic visual world. In contrast to image understanding, video understanding inherently requires modeling temporal dynamics and …

MoVieDrive: Urban Scene Synthesis with Multi-Modal Multi-View Video Diffusion Transformer

2025-08-20 · Guile Wu, David Huang, Dongfeng Bai, Bingbing Liu arxiv

Urban scene synthesis with video generation models has recently shown great potential for autonomous driving. Existing video generation approaches to autonomous driving primarily focus on RGB video generation and lack th…

Scene UnderstandingAutonomous DrivingVideo Generation

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

2026-06-05 · Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao 외 arxiv

Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios. These scenarios require model…