HA-ViD: A Human Assembly Video Dataset for Comprehensive Assembly Knowledge Understanding
Understanding comprehensive assembly knowledge from videos is critical for futuristic ultra-intelligent industry. To enable technological breakthrough, we present HA-ViD - the first human assembly video dataset that features representative industrial assembly scenarios, natural procedural knowledge acquisition process, and consistent human-robot shared annotations. Specifically, HA-ViD captures diverse collaboration patterns of real-world assembly, natural human behaviors and learning progression during assembly, and granulate action annotations to subject, action verb, manipulated object, target object, and tool. We provide 3222 multi-view, multi-modality videos (each video contains one assembly task), 1.5M frames, 96K temporal labels and 2M spatial labels. We benchmark four foundational video understanding tasks: action recognition, action segmentation, object detection and multi-object tracking. Importantly, we analyze their performance for comprehending knowledge in assembly progress, process efficiency, task collaboration, skill parameters and human intention. Details of HA-ViD is available at: https://iai-hrc.github.io/ha-vid.
Code (1)
Tasks
Action RecognitionAction SegmentationMulti-Object TrackingObjectobject-detectionObject DetectionObject TrackingVideo UnderstandingSimilar Papers 제목 키워드 기반
IKEA Manuals at Work: 4D Grounding of Assembly Instructions on Internet Videos
Shape assembly is a ubiquitous task in daily life, integral for constructing complex 3D structures like IKEA furniture. While significant progress has been made in developing autonomous agents for shape assembly, existin…
Pose EstimationSemantic SegmentationVideo Object SegmentationVideo Semantic SegmentationProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly
Assistants on assembly tasks show great potential to benefit humans ranging from helping with everyday tasks to interacting in industrial settings. However, evaluation resources in assembly activities are underexplored. …
A vision-based framework for human behavior understanding in industrial assembly lines
This paper introduces a vision-based framework for capturing and understanding human behavior in industrial assembly lines, focusing on car door manufacturing. The framework leverages advanced computer vision techniques …
Action AnalysisGenerative Timelines for Instructed Visual Assembly
The objective of this work is to manipulate visual timelines (e.g. a video) through natural language instructions, making complex timeline editing tasks accessible to non-expert or potentially even disabled users. We cal…
Language ModellingATTACH Dataset: Annotated Two-Handed Assembly Actions for Human Action Understanding
With the emergence of collaborative robots (cobots), human-robot collaboration in industrial manufacturing is coming into focus. For a cobot to act autonomously and as an assistant, it must understand human actions durin…
Action DetectionAction RecognitionAction UnderstandingVocal Bursts Valence Prediction