paper-with-me

홈 › Papers

A Survey on Safe Multi-Modal Learning System

2024-02-08 · Tianyi Zhao, Liangliang Zhang, Yao Ma, Lu Cheng

In the rapidly evolving landscape of artificial intelligence, multimodal learning systems (MMLS) have gained traction for their ability to process and integrate information from diverse modality inputs. Their expanding use in vital sectors such as healthcare has made safety assurance a critical concern. However, the absence of systematic research into their safety is a significant barrier to progress in this field. To bridge the gap, we present the first taxonomy that systematically categorizes and assesses MMLS safety. This taxonomy is structured around four fundamental pillars that are critical to ensuring the safety of MMLS: robustness, alignment, monitoring, and controllability. Leveraging this taxonomy, we review existing methodologies, benchmarks, and the current state of research, while also pinpointing the principal limitations and gaps in knowledge. Finally, we discuss unique challenges in MMLS safety. In illuminating these challenges, we aim to pave the way for future research, proposing potential directions that could lead to significant advancements in the safety protocols of MMLS.

📄 PDF Abstract BibTeX arXiv:2402.05355

Code (0)

등록된 구현이 없습니다.

Tasks

Survey

Similar Papers 제목 키워드 기반

Safety of Multimodal Large Language Models on Images and Texts

2024-02-01 · Xin Liu, Yichen Zhu, Yunshi Lan, Chao Yang 외

Attracted by the impressive power of Multimodal Large Language Models (MLLMs), the public is increasingly utilizing them to improve the efficiency of daily work. Nonetheless, the vulnerabilities of MLLMs to unsafe instru…

Survey

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks

2026-07-08 · Nobin Sarwar, Shubhashis Roy Dipta, Zheyuan Liu, Vaidehi Patil arxiv

With the growing adoption of VLMs, DMs, LLMs, and AFMs, these multimodal foundation models can inadvertently encode sensitive, copyrighted, biased, or unsafe cross-modal associations that originate from their training da…

Safety in Embodied AI: A Survey of Risks, Attacks, and Defenses

2026-03-28 · Xiao Li, Xiang Zheng, Yifeng Gao, Xinyu Xia 외 arxiv

Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open-world, safety-critical environments. As these systems gain autonomy and enter do…

Jailbreak Attacks and Defenses against Multimodal Generative Models: A Survey

2024-11-14 · Xuannan Liu, Xing Cui, Peipei Li, Zekun Li 외

The rapid evolution of multimodal foundation models has led to significant advancements in cross-modal understanding and generation across diverse modalities, including text, images, audio, and video. However, these mode…

Progressive Bird's Eye View Perception for Safety-Critical Autonomous Driving: A Comprehensive Survey

2025-08-11 · Yan Gong, Naibang Wang, Jianli Lu, Xinyu Zhang 외 arxiv

Bird's-Eye-View (BEV) perception has become a foundational paradigm in autonomous driving, enabling unified spatial representations that support robust multi-sensor fusion and multi-agent collaboration. As autonomous veh…

Autonomous VehiclesAutonomous Driving