paper-with-me

홈 › Papers

SkillSight: Efficient First-Person Skill Assessment with Gaze

2025-11-24 · Chi Hsuan Wu, Kumar Ashutosh, Kristen Grauman arxiv

Egocentric perception on smart glasses could transform how we learn new skills in the physical world, but automatic skill assessment remains a fundamental technical challenge. We introduce SkillSight for power-efficient skill assessment from first-person data. Central to our approach is the hypothesis that skill level is evident not only in how a person performs an activity (video), but also in how they direct their attention when doing so (gaze). Our two-stage framework first learns to jointly model gaze and egocentric video when predicting skill level, then distills a gaze-only student model. At inference, the student model requires only gaze input, drastically reducing power consumption by eliminating continuous video processing. Experiments on three datasets spanning cooking, music, and sports establish, for the first time, the valuable role of gaze in skill understanding across diverse real-world settings. Our SkillSight teacher model achieves state-of-the-art performance, while our gaze-only student variant maintains high accuracy using 73x less power than competing methods. These results pave the way for in-the-wild AI-supported skill learning.

📄 PDF Abstract BibTeX arXiv:2511.19629

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

SkillSight: Calibrating Generic Content Bias for Skill Retrieval

2026-07-21 · Jinying Xiao, Bin Li, Xiaopeng Li, Jianling Li 외 arxiv

As large language model agents gain access to increasingly large skill libraries, retrieving the right skill becomes critical to reliable capability selection and execution. Existing retrievers often treat skill contents…

Are you really looking at me? A Feature-Extraction Framework for Estimating Interpersonal Eye Gaze from Conventional Video

2019-06-21 · Minh Tran, Taylan Sen, Kurtis Haut, Mohammad Rafayet Ali 외

Despite a revolution in the pervasiveness of video cameras in our daily lives, one of the most meaningful forms of nonverbal affective communication, interpersonal eye gaze, i.e. eye gaze relative to a conversation partn…

ClusteringDeception Detection

MOSAIC-F: A Framework for Enhancing Students' Oral Presentation Skills through Personalized Feedback

2025-06-10 · Alvaro Becerra, Daniel Andres, Pablo Villegas, Roberto Daza 외

In this article, we present a novel multimodal feedback framework called MOSAIC-F, an acronym for a data-driven Framework that integrates Multimodal Learning Analytics (MMLA), Observations, Sensors, Artificial Intelligen…

GazeLLM: Multimodal LLMs incorporating Human Visual Attention

2025-03-31 · Jun Rekimoto

Large Language Models (LLMs) are advancing into Multimodal LLMs (MLLMs), capable of processing image, audio, and video as well as text. Combining first-person video, MLLMs show promising potential for understanding human…

NeuGaze: Reshaping the future BCI

2025-04-21 · Yiqian Yang

Traditional brain-computer interfaces (BCIs), reliant on costly electroencephalography or invasive implants, struggle with complex human-computer interactions due to setup complexity and limited precision. We present Neu…