paper-with-me

Papers

From Videos to Conversations: Egocentric Instructions for Task Assistance

2026-02-01 · Lavisha Aggarwal, Vikas Bahirwani, Andrea Colaco arxiv

Many everyday tasks, ranging from appliance repair and cooking to car maintenance, require expert knowledge, particularly for complex, multi-step procedures. Despite growing interest in AI agents for augmented reality (AR) assistance, progress remains limited by the scarcity of large-scale multimodal conversational datasets grounded in real-world task execution, in part due to the cost and logistical complexity of human-assisted data collection. In this paper, we present a framework to automatically transform single person instructional videos into two-person multimodal task-guidance conversations. Our fully automatic pipeline, based on large language models, provides a scalable and cost efficient alternative to traditional data collection approaches. Using this framework, we introduce HowToDIV, a multimodal dataset comprising 507 conversations, 6,636 question answer pairs, and 24 hours of video spanning multiple domains. Each session consists of a multi-turn expert-novice interaction. Finally, we report baseline results using Gemma 3 and Qwen 2.5 on HowToDIV, providing an initial benchmark for multimodal procedural task assistance.

📄 PDF Abstract BibTeX arXiv:2602.01038

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

2024-12-30 · Yifei HUANG, Jilan Xu, Baoqi Pei, Yuping He 외

We introduce Vinci, a real-time embodied smart assistant built upon an egocentric vision-language model. Designed for deployment on portable devices such as smartphones and wearable cameras, Vinci operates in an "always …

Language ModelingLanguage ModellingTask PlanningVideo Generation

Generating Dialogues from Egocentric Instructional Videos for Task Assistance: Dataset, Method and Benchmark

2025-08-15 · Lavisha Aggarwal, Vikas Bahirwani, Lin Li, Andrea Colaco arxiv

Many everyday tasks ranging from fixing appliances, cooking recipes to car maintenance require expert knowledge, especially when tasks are complex and multi-step. Despite growing interest in AI agents, there is a scarcit…

EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering

2025-02-11 · CVPR 2025 1 · Sheng Zhou, Junbin Xiao, Qingyun Li, Yicong Li 외

We introduce EgoTextVQA, a novel and rigorously constructed benchmark for egocentric QA assistance involving scene text. EgoTextVQA contains 1.5K ego-view videos and 7K scene-text aware questions that reflect real user n…

Question AnsweringVideo Question Answering

VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios

2026-05-11 · Yu-Hsiang Liu, Yu-Chien Tang, An-Zi Yen arxiv

Evaluating whether AI agents can proactively assist humans in daily activities, ranging from routine household tasks to urgent safety-critical situations, requires diverse visual data. However, collecting such scenarios …

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

2026-07-13 · Gong Sitong, Tianyu Yan, Caixin Kang, Bo Zheng 외 hf

When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive. Yet e…