paper-with-me

홈 › Papers

Vid2Coach: Transforming How-To Videos into Task Assistants

2025-05-31 · Mina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh, Kristen Grauman, Amy Pavel

People use videos to learn new recipes, exercises, and crafts. Such videos remain difficult for blind and low vision (BLV) people to follow as they rely on visual comparison. Our observations of visual rehabilitation therapists (VRTs) guiding BLV people to follow how-to videos revealed that VRTs provide both proactive and responsive support including detailed descriptions, non-visual workarounds, and progress feedback. We propose Vid2Coach, a system that transforms how-to videos into wearable camera-based assistants that provide accessible instructions and mixed-initiative feedback. From the video, Vid2Coach generates accessible instructions by augmenting narrated instructions with demonstration details and completion criteria for each step. It then uses retrieval-augmented-generation to extract relevant non-visual workarounds from BLV-specific resources. Vid2Coach then monitors user progress with a camera embedded in commercial smart glasses to provide context-aware instructions, proactive feedback, and answers to user questions. BLV participants (N=8) using Vid2Coach completed cooking tasks with 58.5\% fewer errors than when using their typical workflow and wanted to use Vid2Coach in their daily lives. Vid2Coach demonstrates an opportunity for AI visual assistance that strengthens rather than replaces non-visual expertise.

📄 PDF Abstract BibTeX arXiv:2506.00717

Code (0)

등록된 구현이 없습니다.

Tasks

Retrieval-augmented Generation

Similar Papers 제목 키워드 기반

Generative AI in Training and Coaching: Redefining the Design Process of Learning Materials

2025-08-06 · Alexander Komar, Marc-André Heidelmann, Kristina Schaaff arxiv

Generative artificial intelligence (GenAI) is transforming education, redefining the role of trainers and coaches in learning environments. In our study, we explore how AI integrates into the design process of learning m…

Who Owns the Text? Design Patterns for Preserving Authorship in AI-Assisted Writing

2026-01-15 · Bohan Zhang, Chengke Bu, Paramveer S. Dhillon arxiv

AI writing assistants can reduce effort and improve fluency, but they may also weaken writers' sense of authorship. We study this tension with an ownership-aware co-writing editor that offers on-demand, sentence-level su…

PresentCoach: Dual-Agent Presentation Coaching through Exemplars and Interactive Feedback

2025-11-19 · Sirui Chen, Jinsong Zhou, Xinli Xu, Xiaoyu Yang 외 arxiv

Effective presentation skills are essential in education, professional communication, and public speaking, yet learners often lack access to high-quality exemplars or personalized coaching. Existing AI tools typically pr…

Streaming Interventions: Can Video Large Language Models Correct Mistakes as They Occur?

2026-06-08 · Apratim Bhattacharyya, Shweta Mahajan, Sanjay Haresh, Rajeev Yasarla 외 arxiv

Learning everyday skills, like cooking a dish, relies increasingly on instructional media such as online videos. This opens the door to the use of video (and multimodal) large language models (LLMs) as task guidance assi…

Behavior-Adaptive Conversational Agents: Toward a Fluid Personality Framework

2026-07-01 · Hasibur Rahman, Smit Desai arxiv

Large language model (LLM)-based conversational agents (CAs) are now ubiquitous, creating new opportunities for AI-mediated behavior change. Their capacity to project nuanced personalities and adopt diverse metaphorical …