paper-with-me

Papers

Can Vision-Language Models Understand Construction Workers? An Exploratory Study

2026-01-15 · Hieu Bui, Nathaniel E. Chodosh, Arash Tavakoli arxiv

As robotics become increasingly integrated into construction workflows, their ability to interpret and respond to human behavior will be essential for enabling safe and effective collaboration. Vision-Language Models (VLMs) have emerged as a promising tool for visual understanding tasks and offer the potential to recognize human behaviors without extensive domain-specific training. This capability makes them particularly appealing in the construction domain, where labeled data is scarce and monitoring worker actions and emotional states is critical for safety and productivity. In this study, we evaluate the performance of three leading VLMs, GPT-4o, Florence 2, and LLaVa-1.5, in detecting construction worker actions and emotions from static site images. Using a curated dataset of 1,000 images annotated across ten action and ten emotion categories, we assess each model's outputs through standardized inference pipelines and multiple evaluation metrics. GPT-4o consistently achieved the highest scores across both tasks, with an average F1-score of 0.756 and accuracy of 0.799 in action recognition, and an F1-score of 0.712 and accuracy of 0.773 in emotion recognition. Florence 2 performed moderately, with F1-scores of 0.497 for action and 0.414 for emotion, while LLaVa-1.5 showed the lowest overall performance, with F1-scores of 0.466 for action and 0.461 for emotion. Confusion matrix analyses revealed that all models struggled to distinguish semantically close categories, such as collaborating in teams versus communicating with supervisors. While the results indicate that general-purpose VLMs can offer a baseline capability for human behavior recognition in construction environments, further improvements, such as domain adaptation, temporal modeling, or multimodal sensing, may be needed for real-world reliability.

📄 PDF Abstract BibTeX arXiv:2601.10835

Code (0)

등록된 구현이 없습니다.

Tasks

Emotion RecognitionAction RecognitionDomain Adaptation

Similar Papers 제목 키워드 기반

Automating Analysis of Construction Workers Viewing Patterns for Personalized Safety Training and Management

2018-08-20 · Jeelani Idris, Han Kevin, Albert Alex

Unrecognized hazards increase the likelihood of workplace fatalities and injuries substantially. However, recent research has demonstrated that a large proportion of hazards remain unrecognized in dynamic construction en…

Management

Natural Language Instructions for Intuitive Human Interaction with Robotic Assistants in Field Construction Work

2023-07-09 · Somin Park, Xi Wang, Carol C. Menassa, Vineet R. Kamat 외

The introduction of robots is widely considered to have significant potential of alleviating the issues of worker shortage and stagnant productivity that afflict the construction industry. However, it is challenging to u…

Language ModellingNatural Language UnderstandingTAG

Real-world Mapping of Gaze Fixations Using Instance Segmentation for Road Construction Safety Applications

2019-01-30 · Idris Jeelani, Khashayar Asadi, Hariharan Ramshankar, Kevin Han 외

Research studies have shown that a large proportion of hazards remain unrecognized, which expose construction workers to unanticipated safety risks. Recent studies have also found that a strong correlation exists between…

Instance SegmentationSemantic SegmentationTransfer Learning

ErgoChat: a Visual Query System for the Ergonomic Risk Assessment of Construction Workers

2024-12-27 · Chao Fan, Qipei Mei, Xiaonan Wang, Xinming Li

In the construction sector, workers often endure prolonged periods of high-intensity physical work and prolonged use of tools, resulting in injuries and illnesses primarily linked to postural ergonomic risks, a longstand…

Image CaptioningQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Integration of Object Detection and Small VLMs for Construction Safety Hazard Identification

2026-04-06 · Muhammad Adil, Mehmood Ahmed, Muhammad Aqib, Vicente A. Gonzalez 외 arxiv

Accurate and timely identification of construction hazards around workers is essential for preventing workplace accidents. While large vision-language models (VLMs) demonstrate strong contextual reasoning capabilities, t…

Multimodal ReasoningObject Detection