paper-with-me

Papers

ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly

2025-09-03 · Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada, Susan Holm, Yuran Wang, Vincent Zhou, Ken Fukuda, Teruko Mitamura arxiv

Assistants on assembly tasks show great potential to benefit humans ranging from helping with everyday tasks to interacting in industrial settings. However, evaluation resources in assembly activities are underexplored. To foster system development, we propose a new multimodal QA evaluation dataset on assembly activities. Our dataset, ProMQA-Assembly, consists of 646 QA pairs that require multimodal understanding of human activity videos and their instruction manuals in an online-style manner. For cost effectiveness in the data creation, we adopt a semi-automated QA annotation approach, where LLMs generate candidate QA pairs and humans verify them. We further improve QA generation by integrating fine-grained action labels to diversify question types. Additionally, we create 81 instruction task graphs for our target assembly tasks. These newly created task graphs are used in our benchmarking experiment, as well as in facilitating the human verification process. With our dataset, we benchmark models, including competitive proprietary multimodal models. We find that ProMQA-Assembly contains challenging multimodal questions, where reasoning models showcase promising results. We believe our new evaluation dataset contributes to the further development of procedural-activity assistants.

📄 PDF Abstract BibTeX arXiv:2509.02949

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TAMA: Tool-Augmented Multimodal Agent for Procedural Activity Understanding

2025-09-30 · Kimihiro Hasegawa, Wiradee Imrattanatrai, Masaki Asada, Ken Fukuda 외 arxiv

Procedural activity assistants potentially support humans in a variety of settings, from our daily lives, e.g., cooking or assembling flat-pack furniture, to professional situations, e.g., manufacturing or biological exp…

Multimodal Reasoning

ProMQA: Question Answering Dataset for Multimodal Procedural Activity Understanding

2024-10-29 · Kimihiro Hasegawa, Wiradee Imrattanatrai, Zhi-Qi Cheng, Masaki Asada 외

Multimodal systems have great potential to assist humans in procedural activities, where people follow instructions to achieve their goals. Despite diverse application scenarios, systems are typically evaluated on tradit…

Action RecognitionAction SegmentationQuestion AnsweringTemporal Action Segmentation

LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants

2025-07-07 · Haochen Huang, Jiahuan Pei, Mohammad Aliannejadi, Xin Sun 외 arxiv

Vision-language models (VLMs) are facing the challenges of understanding and following multimodal assembly instructions, particularly when fine-grained spatial reasoning and precise object state detection are required. I…

Spatial ReasoningObject Detection

IndEgo: A Dataset of Industrial Scenarios and Collaborative Work for Egocentric Assistants

2025-11-24 · Vivek Chavan, Yasmina Imgrund, Tung Dao, Sanwantri Bai 외 arxiv

We introduce IndEgo, a multimodal egocentric and exocentric dataset addressing common industrial tasks, including assembly/disassembly, logistics and organisation, inspection and repair, woodworking, and others. The data…

Question Answering

IMPACT: A Dataset for Multi-Granularity Human Procedural Action Understanding in Industrial Assembly

2026-04-12 · Di Wen, Zeyun Zhong, David Schneider, Manuel Zaremski 외 arxiv

We introduce IMPACT, a synchronized five-view RGB-D dataset for deployment-oriented industrial procedural understanding, built around real assembly and disassembly of a commercial angle grinder with professional-grade to…

Action Understanding