paper-with-me

Papers

ExpertAF: Expert Actionable Feedback from Video

2024-08-01 · CVPR 2025 1 · Kumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos, Kris Kitani, Kristen Grauman

Feedback is essential for learning a new skill or improving one's current skill-level. However, current methods for skill-assessment from video only provide scores or compare demonstrations, leaving the burden of knowing what to do differently on the user. We introduce a novel method to generate actionable feedback from video of a person doing a physical activity, such as basketball or soccer. Our method takes a video demonstration and its accompanying 3D body pose and generates (1) free-form expert commentary describing what the person is doing well and what they could improve, and (2) a visual expert demonstration that incorporates the required corrections. We show how to leverage Ego-Exo4D's videos of skilled activity and expert commentary together with a strong language model to create a weakly-supervised training dataset for this task, and we devise a multimodal video-language model to infer coaching feedback. Our method is able to reason across multi-modal input combinations to output full-spectrum, actionable coaching -- expert commentary, expert video retrieval, and expert pose generation -- outperforming strong vision-language models on both established metrics and human preference studies. Code and data will be publicly released.

📄 PDF Abstract BibTeX arXiv:2408.00672

Code (0)

등록된 구현이 없습니다.

Tasks

Language ModelingLanguage ModellingVideo Retrieval

Similar Papers 제목 키워드 기반

AIDE: Automated Instruction via Distilled Expertise for Reference-Free Motor Skill Coaching

2026-08-04 · Yoshiki Ito arxiv

Generating natural-language coaching feedback on motor skills can accelerate learning, yet expert coaches are scarce and expensive. Existing reference-based methods require expert demonstrations at both training and infe…

Learning Consistent Temporal Grounding between Related Tasks in Sports Coaching

2026-03-19 · Arushi Rai, Adriana Kovashka arxiv

Video-LLMs often attend to irrelevant frames, which is especially detrimental for sports coaching tasks requiring precise temporal grounding. Yet obtaining frame-level supervision is challenging: expensive to collect fro…

Explainable AI for Automated User-specific Feedback in Surgical Skill Acquisition

2025-08-04 · Catalina Gomez, Lalithkumar Seenivasan, Xinrui Zou, Jeewoo Yoon 외 arxiv

Traditional surgical skill acquisition relies heavily on expert feedback, yet direct access is limited by faculty availability and variability in subjective assessments. While trainees can practice independently, the lac…

Learning Skill-Attributes for Transferable Assessment in Video

2025-11-17 · Kumar Ashutosh, Kristen Grauman arxiv

Skill assessment from video entails rating the quality of a person's physical performance and explaining what could be done better. Today's models specialize for an individual sport, and suffer from the high cost and sca…

ProfVLM: A lightweight video-language model for multi-view proficiency estimation

2025-09-30 · Edoardo Bianchi, Jacopo Staiano, Antonio Liotta arxiv

Most existing approaches formulate action quality assessment and skill proficiency estimation as discriminative prediction tasks, typically producing discrete labels or scores without explicitly modeling the reasoning pr…

Action Quality Assessment