paper-with-me

홈 › Papers

HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection

2025-09-28 · Siyuan Gao, Jiashu Yao, Haoyu Wen, Yuhang Guo, Zeming Liu, Heyan Huang arxiv

Safety hazards in the home are a leading cause of preventable domestic injuries, motivating an automated inspector that actively explores a home and reports hazards before they cause harm. We introduce HomeSafeBench, the first benchmark for free-exploration home safety inspection with egocentric visual feedback, in which an embodied agent navigates a fully interactive 3D home, adjusts its viewpoint, and reports hazards purely from rendered first-person views. Built on the VirtualHome simulator, it covers five categories of common household hazards and comprises 1,000 human-validated inspection tasks. Evaluating a broad range of state-of-the-art Vision-Language Models (VLMs) reveals a large gap, where the best model reaches only about 34.7% F1, far below the 98.0% of a human inspector. Moreover, precision far exceeds recall across models, revealing a systematic tendency to under-report hazards that reflects a shared deficiency in risk recognition. To close this gap at low cost, we propose CueBack, an offline data-construction method that exploits the clue-precedes-confirmation structure of inspection, backtracking a privileged trajectory to the earliest frame where a hazard cue becomes visible and rewriting it into executable supervision. Fine-tuning a 4B-size VLM on CueBack-constructed data raises the average F1 from 18.7% to 45.3% on an out-of-distribution test set, surpassing the strongest closed-source model performance 34.7%. The benchmark, training dataset, and code are available at https://github.com/BITHLP/HomeSafeBench.

📄 PDF Abstract BibTeX arXiv:2509.23690

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning

2025-10-13 · Ganlin Yang, Tianyi Zhang, Haoran Hao, Weiyun Wang 외 arxiv

While significant research has focused on developing embodied reasoning capabilities using Vision-Language Models (VLMs) or integrating advanced VLMs into Vision-Language-Action (VLA) models for end-to-end robot control,…

Spatial Reasoning

FreeAskWorld: An Interactive and Closed-Loop Simulator for Human-Centric Embodied AI

2025-11-17 · Yuhang Peng, Yizhou Pan, Xinning He, Jihaoyu Yang 외 arxiv

As embodied intelligence emerges as a core frontier in artificial intelligence research, simulation platforms must evolve beyond low-level physical interactions to capture complex, human-centered social behaviors. We int…

VLMbench: A Compositional Benchmark for Vision-and-Language Manipulation

2022-06-17 · Kaizhi Zheng, Xiaotong Chen, Odest Chadwicke Jenkins, Xin Eric Wang

Benefiting from language flexibility and compositionality, humans naturally intend to use language to command an embodied agent for complex tasks such as navigation and object manipulation. In this work, we aim to fill t…

Object

EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models

2024-06-09 · Mengfei Du, Binhao Wu, Zejun Li, Xuanjing Huang 외

The recent rapid development of Large Vision-Language Models (LVLMs) has indicated their potential for embodied tasks.However, the critical skill of spatial understanding in embodied environments has not been thoroughly …

Benchmarking

GenRL: Multimodal-foundation world models for generalization in embodied agents

2024-06-26 · Pietro Mazzaglia, Tim Verbelen, Bart Dhoedt, Aaron Courville 외

Learning generalist embodied agents, able to solve multitudes of tasks in different domains is a long-standing problem. Reinforcement learning (RL) is hard to scale up as it requires a complex reward design for each task…

BenchmarkingReinforcement Learning (RL)