paper-with-me

홈 › Papers

When Helpers Become Hazards: A Benchmark for Analyzing Multimodal LLM-Powered Safety in Daily Life

2026-01-07 · Xinyue Lou, Jinan Xu, Jingyi Yin, Xiaolong Wang, Zhaolu Kang, Youwei Liao, Yixuan Wang, Xiangyu Shi, Fengran Mo, Su Yao, Kaiyu Huang arxiv

As Multimodal Large Language Models (MLLMs) become an indispensable assistant in human life, the unsafe content generated by MLLMs poses a danger to human behavior, perpetually overhanging human society like a sword of Damocles. To investigate and evaluate the safety impact of MLLMs responses on human behavior in daily life, we introduce SaLAD, a multimodal safety benchmark which contains 2,013 real-world image-text samples across 10 common categories, with a balanced design covering both unsafe scenarios and cases of oversensitivity. It emphasizes realistic risk exposure, authentic visual inputs, and fine-grained cross-modal reasoning, ensuring that safety risks cannot be inferred from text alone. We further propose a safety-warning-based evaluation framework that encourages models to provide clear and informative safety warnings, rather than generic refusals. Results on 18 MLLMs demonstrate that the top-performing models achieve a safe response rate of only 57.2% on unsafe queries. Moreover, even popular safety alignment methods limit effectiveness of the models in our scenario, revealing the vulnerabilities of current MLLMs in identifying dangerous behaviors in daily life. Our dataset is available at https://github.com/xinyuelou/SaLAD.

📄 PDF Abstract BibTeX arXiv:2601.04043

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Online library learning in human visual puzzle solving

2026-03-24 · Pinzhe Zhao, Emanuele Sansone, Marta Kryven, Bonan Zhao arxiv

When learning a novel complex task, people often form efficient reusable abstractions that simplify future work, despite uncertainty about the future. We study this process in a visual puzzle task where participants defi…

HomeSafeBench: A Benchmark for Embodied Vision-Language Models in Free-Exploration Home Safety Inspection

2025-09-28 · Siyuan Gao, Jiashu Yao, Haoyu Wen, Yuhang Guo 외 arxiv

Safety hazards in the home are a leading cause of preventable domestic injuries, motivating an automated inspector that actively explores a home and reports hazards before they cause harm. We introduce HomeSafeBench, the…

Learning from data in the mixed adversarial non-adversarial case: Finding the helpers and ignoring the trolls

2022-08-05 · Da Ju, Jing Xu, Y-Lan Boureau, Jason Weston

The promise of interaction between intelligent conversational agents and humans is that models can learn from such feedback in order to improve. Unfortunately, such exchanges in the wild will not always involve human utt…

Efficient and Debiased Learning of Average Hazard Under Non-Proportional Hazards

2026-02-13 · Xiang Meng, Lu Tian, Kenneth Kehl, Hajime Uno arxiv

The hazard ratio from the Cox proportional hazards model is a ubiquitous summary of treatment effect. However, when hazards are non-proportional, the hazard ratio can lose a stable causal interpretation and become study-…

R2H: Building Multimodal Navigation Helpers that Respond to Help Requests

2023-05-23 · Yue Fan, Jing Gu, Kaizhi Zheng, Xin Eric Wang

Intelligent navigation-helper agents are critical as they can navigate users in unknown areas through environmental awareness and conversational ability, serving as potential accessibility tools for individuals with disa…

BenchmarkingLanguage ModelingLanguage ModellingLarge Language Model+2