paper-with-me

홈 › Papers

Evaluating Cascaded Methods of Vision-Language Models for Zero-Shot Detection and Association of Hardhats for Increased Construction Safety

2024-10-16 · Lucas Choi, Ross Greer

This paper evaluates the use of vision-language models (VLMs) for zero-shot detection and association of hardhats to enhance construction safety. Given the significant risk of head injuries in construction, proper enforcement of hardhat use is critical. We investigate the applicability of foundation models, specifically OWLv2, for detecting hardhats in real-world construction site images. Our contributions include the creation of a new benchmark dataset, Hardhat Safety Detection Dataset, by filtering and combining existing datasets and the development of a cascaded detection approach. Experimental results on 5,210 images demonstrate that the OWLv2 model achieves an average precision of 0.6493 for hardhat detection. We further analyze the limitations and potential improvements for real-world applications, highlighting the strengths and weaknesses of current foundation models in safety perception domains.

📄 PDF Abstract BibTeX arXiv:2410.12225

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation

2024-06-02 · Yunheng Li, Zhongyu Li, Quansheng Zeng, Qibin Hou 외

Pre-trained vision-language models, e.g., CLIP, have been successfully applied to zero-shot semantic segmentation. Existing CLIP-based approaches primarily utilize visual features from the last layer to align with text e…

SegmentationSemantic SegmentationZero-Shot Semantic Segmentation

Evaluating Vision-Language Models for Zero-Shot Detection, Classification, and Association of Motorcycles, Passengers, and Helmets

2024-08-05 · Lucas Choi, Ross Greer

Motorcycle accidents pose significant risks, particularly when riders and passengers do not wear helmets. This study evaluates the efficacy of an advanced vision-language foundation model, OWLv2, in detecting and classif…

Zero-Shot Learning

Enhancing Fine-Grained Image Classifications via Cascaded Vision Language Models

2024-05-18 · Canshi Wei

Fine-grained image classification, particularly in zero/few-shot scenarios, presents a significant challenge for vision-language models (VLMs), such as CLIP. These models often struggle with the nuanced task of distingui…

Fine-Grained Image Classificationimage-classificationImage Classification

Cerberus: Real-Time Video Anomaly Detection via Cascaded Vision-Language Models

2025-10-18 · Yue Zheng, Xiufang Shi, Jiming Chen, Yuanchao Shu arxiv

Video anomaly detection (VAD) has rapidly advanced by recent development of Vision-Language Models (VLMs). While these models offer superior zero-shot detection capabilities, their immense computational cost and unstable…

Video Anomaly DetectionVisual Grounding

SageLM: A Multi-aspect and Explainable Large Language Model for Speech Judgement

2025-08-28 · Yuan Ge, Junxiang Zhang, Xiaoqian Liu, Bei Li 외 arxiv

Speech-to-Speech (S2S) Large Language Models (LLMs) are foundational to natural human-computer interaction, enabling end-to-end spoken dialogue systems. However, evaluating these models remains a fundamental challenge. W…

Reinforcement Learning