paper-with-me

Papers

The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning

2026-01-20 · Renmiao Chen, Yida Lu, Shiyao Cui, Xuan Ouyang, Victor Shea-Jay Huang, Shumin Zhang, Chengwei Pan, Han Qiu, Minlie Huang arxiv

As Multimodal Large Language Models (MLLMs) acquire stronger reasoning capabilities to handle complex, multi-image instructions, this advancement may pose new safety risks. We study this problem by introducing MIR-SafetyBench, the first benchmark focused on multi-image reasoning safety, which consists of 2,676 instances across a taxonomy of 9 multi-image relations. Our extensive evaluations on 19 MLLMs reveal a troubling trend: models with more advanced multi-image reasoning can be more vulnerable on MIR-SafetyBench. Beyond attack success rates, we find that many responses labeled as safe are superficial, often driven by misunderstanding or evasive, non-committal replies. We further observe that unsafe generations exhibit lower attention entropy than safe ones on average. This internal signature suggests a possible risk that models may over-focus on task solving while neglecting safety constraints. Our code and data are available at https://github.com/thu-coai/MIR-SafetyBench.

📄 PDF Abstract BibTeX arXiv:2601.14127

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Algorithmic decision-making in AVs: Understanding ethical and technical concerns for smart cities

2019-10-29 · Hazel Si Min Lim, Araz Taeihagh

Autonomous Vehicles (AVs) are increasingly embraced around the world to advance smart mobility and more broadly, smart, and sustainable cities. Algorithms form the basis of decision-making in AVs, allowing them to perfor…

Autonomous VehiclesDecision MakingEthics

Being Accountable is Smart: Navigating the Technical and Regulatory Landscape of AI-based Services for Power Grid

2024-08-02 · Anna Volkova, Mahdieh Hatamian, Alina Anapyanova, Hermann de Meer

The emergence of artificial intelligence and digitization of the power grid introduced numerous effective application scenarios for AI-based services for the smart grid. Nevertheless, adopting AI in critical infrastructu…

AI Research Considerations for Human Existential Safety (ARCHES)

2020-05-30 · Andrew Critch, David Krueger

Framed in positive terms, this report examines how technical AI research might be steered in a manner that is more attentive to humanity's long-term prospects for survival as a species. In negative terms, we ask what exi…

Autonomous Alignment with Human Value on Altruism through Considerate Self-imagination and Theory of Mind

2024-12-31 · Haibo Tong, Enmeng Lu, Yinqian Sun, Zhengqiang Han 외

With the widespread application of Artificial Intelligence (AI) in human society, enabling AI to autonomously align with human values has become a pressing issue to ensure its sustainable development and benefit to human…

Benchmarking Safety Risks of Knowledge-Intensive Reasoning under Malicious Knowledge Editing

2026-05-11 · Qinghua Mao, Xi Lin, Jinze Gu, Jun Wu 외 arxiv

Large language models (LLMs) increasingly rely on knowledge editing to support knowledge-intensive reasoning, but this flexibility also introduces critical safety risks: adversaries can inject malicious or misleading kno…

knowledge editing