paper-with-me

홈 › Papers

Parameter-Efficient Multi-View Proficiency Estimation: From Discriminative Classification to Generative Feedback

2026-05-05 · Edoardo Bianchi, Antonio Liotta arxiv

Estimating how well a person performs an action, rather than which action is performed, is central to coaching, rehabilitation, and talent identification. This task is challenging because proficiency is encoded in subtle differences in timing, balance, body mechanics, and execution, often distributed across multiple views and short temporal events. We discuss three recent contributions to multi-view proficiency estimation on Ego-Exo4D. SkillFormer introduces a parameter-efficient discriminative architecture for selective multi-view fusion; PATS improves temporal sampling by preserving locally dense excerpts of fundamental movements; and ProfVLM reformulates proficiency estimation as conditional language generation, producing both a proficiency label and expert-style feedback through a gated cross-view projector and a compact language backbone. Together, these methods achieve state-of-the-art accuracy on Ego-Exo4D with up to 20x fewer trainable parameters and up to 3x fewer training epochs than video-transformer baselines, while moving from closed-set classification toward interpretable feedback generation. These results highlight a shift toward efficient, multi-view systems that combine selective fusion, proficiency-aware sampling, and actionable generative feedback.

📄 PDF Abstract BibTeX arXiv:2605.03848

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ProfVLM: A lightweight video-language model for multi-view proficiency estimation

2025-09-30 · Edoardo Bianchi, Jacopo Staiano, Antonio Liotta arxiv

Most existing approaches formulate action quality assessment and skill proficiency estimation as discriminative prediction tasks, typically producing discrete labels or scores without explicitly modeling the reasoning pr…

Action Quality Assessment

SkillFormer: Unified Multi-View Video Understanding for Proficiency Estimation

2025-05-13 · Edoardo Bianchi, Antonio Liotta

Assessing human skill levels in complex activities is a challenging problem with applications in sports, rehabilitation, and training. In this work, we present SkillFormer, a parameter-efficient architecture for unified …

Computational EfficiencyVideo Understanding

Moving Beyond More Views: Redundancy-Aware Ego-Exo Fusion for Proficiency Estimation

2026-08-26 · Xu Dong, Wanqing Li, Anthony Adeyemi-Ejeye, Andrew Gilbert arxiv

EgoExo proficiency estimation aims to assess action quality by integrating fine-grained motion cues from egocentric (1st-person) views with spatial context from multiple exocentric (3rd-person) views. Simply adding more …

CuriosAI Submission to the EgoExo4D Proficiency Estimation Challenge 2025

2025-07-08 · Hayato Tanoue, Hiroki Nishihara, Yuma Suzuki, Takayuki Hori 외 arxiv

This report presents the CuriosAI team's submission to the EgoExo4D Proficiency Estimation Challenge at CVPR 2025. We propose two methods for multi-view skill assessment: (1) a multi-task learning framework using Sapiens…

Multi-Task Learning

SkillMoV: Mixture-of-View Routing with Prototype-Conditioned Gating for Unified Multi-View Proficiency Estimation

2026-06-16 · Edoardo Bianchi, Antonio Liotta arxiv

Estimating human proficiency from video is a key challenge for automated skill assessment, with applications in sports coaching, music pedagogy, surgical training, and workplace learning. Existing approaches often focus …