paper-with-me

홈 › Papers

ProSkill: Segment-Level Skill Assessment in Procedural Videos

2026-01-28 · Michele Mazzamuto, Daniele Di Mauro, Gianpiero Francesca, Giovanni Maria Farinella, Antonino Furnari arxiv

Skill assessment in procedural videos is crucial for the objective evaluation of human performance in settings such as manufacturing and procedural daily tasks. Current research on skill assessment has predominantly focused on sports and lacks large-scale datasets for complex procedural activities. Existing studies typically involve only a limited number of actions, focus on either pairwise assessments (e.g., A is better than B) or on binary labels (e.g., good execution vs needs improvement). In response to these shortcomings, we introduce ProSkill, the first benchmark dataset for action-level skill assessment in procedural tasks. ProSkill provides absolute skill assessment annotations, along with pairwise ones. This is enabled by a novel and scalable annotation protocol that allows for the creation of an absolute skill assessment ranking starting from pairwise assessments. This protocol leverages a Swiss Tournament scheme for efficient pairwise comparisons, which are then aggregated into consistent, continuous global scores using an ELO-based rating system. We use our dataset to benchmark the main state-of-the-art skill assessment algorithms, including both ranking-based and pairwise paradigms. The suboptimal results achieved by the current state-of-the-art highlight the challenges and thus the value of ProSkill in the context of skill assessment for procedural videos. All data and code are available at https://fpv-iplab.github.io/ProSkill/

📄 PDF Abstract BibTeX arXiv:2601.20661

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Automated Procedural Analysis via Video-Language Models for AI-assisted Nursing Skills Assessment

2025-09-20 · Shen Chang, Dennis Liu, Renran Tian, Kristen L. Swartzell 외 arxiv

Consistent high-quality nursing care is essential for patient safety, yet current nursing education depends on subjective, time-intensive instructor feedback in training future nurses, which limits scalability and effici…

Action RecognitionSkills Assessment

SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos

2026-04-10 · Xiyang Huang, Jiawei Lin, Keying Wu, Jiaxin Huang 외 arxiv

Current video benchmarks for multimodal large language models (MLLMs) focus on event recognition, temporal ordering, and long-context recall, but overlook a harder capability required for expert procedural judgment: trac…

SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries

2026-08-06 · Xingyu Tan, Xiaoyang Wang, Qing Liu, Xiwei Xu 외 arxiv

Large Language Models (LLMs) increasingly act as agents whose procedural knowledge is stored in reusable skill packages and loaded at inference time. As skill libraries grow, a central challenge is to expose the smallest…

Anything2Skill: Compiling External Knowledge into Reusable Skills for Agents

2026-06-08 · Qianjun Pan, Yutao Yang, Junsong Li, Jie Zhou 외 arxiv

Retrieval-augmented generation (RAG) enables agents to access external knowledge at inference time, but it primarily retrieves fragmented declarative evidence, leaving agents to repeatedly infer task procedures from pass…

Deep Learning with Convolutional Neural Network for Objective Skill Evaluation in Robot-assisted Surgery

2018-06-15 · Ziheng Wang, Ann Majewicz Fey

With the advent of robot-assisted surgery, the role of data-driven approaches to integrate statistics and machine learning is growing rapidly with prominent interests in objective surgical skill assessment. However, most…

Time SeriesTime Series Analysis