paper-with-me

Papers

Pushing the Limits of Safety: A Technical Report on the ATLAS Challenge 2025

2025-06-14 · Zonghao Ying, Siyang Wu, Run Hao, Peng Ying, Shixuan Sun, PengYu Chen, Junze Chen, Hao Du, Kaiwen Shen, Shangkun Wu, Jiwei Wei, Shiyuan He, Yang Yang, Xiaohai Xu, Ke Ma, Qianqian Xu, Qingming Huang, Shi Lin, Xun Wang, Changting Lin, Meng Han, Yilei Jiang, Siqi Lai, Yaozhi Zheng, Yifei Song, Xiangyu Yue, Zonglei Jing, Tianyuan Zhang, Zhilei Zhu, Aishan Liu, Jiakai Wang, Siyuan Liang, Xianglong Kong, Hainan Li, Junjie Mu, Haotong Qin, Yue Yu, Lei Chen, Felix Juefei-Xu, Qing Guo, Xinyun Chen, Yew Soon Ong, Xianglong Liu, Dawn Song, Alan Yuille, Philip Torr, DaCheng Tao

Multimodal Large Language Models (MLLMs) have enabled transformative advancements across diverse applications but remain susceptible to safety threats, especially jailbreak attacks that induce harmful outputs. To systematically evaluate and improve their safety, we organized the Adversarial Testing & Large-model Alignment Safety Grand Challenge (ATLAS) 2025}. This technical report presents findings from the competition, which involved 86 teams testing MLLM vulnerabilities via adversarial image-text attacks in two phases: white-box and black-box evaluations. The competition results highlight ongoing challenges in securing MLLMs and provide valuable guidance for developing stronger defense mechanisms. The challenge establishes new benchmarks for MLLM safety evaluation and lays groundwork for advancing safer multimodal AI systems. The code and data for this challenge are openly available at https://github.com/NY1024/ATLAS_Challenge_2025.

📄 PDF Abstract BibTeX arXiv:2506.12430

Code (1)

ny1024/atlas_challenge_2025 공식 구현

Similar Papers 제목 키워드 기반

Embedding Atlas: Low-Friction, Interactive Embedding Visualization

2025-05-09 · Donghao Ren, Fred Hohman, Halden Lin, Dominik Moritz

Embedding projections are popular for visualizing large datasets and models. However, people often encounter "friction" when using embedding visualization tools: (1) barriers to adoption, e.g., tedious data wrangling and…

Friction

(hu)Man vs. Machine: In the Future of Motorsport, can Autonomous Vehicles Compete?

2026-03-02 · Armand Amaritei, Amber-Lily Blackadder, Sebastian Donnelly, Lora Hernandez 외 arxiv

Motorsport has historically driven technological innovation in the automotive industry. Autonomous racing provides a proving ground to push the limits of performance of autonomous vehicle (AV) systems. In principle, AVs …

Autonomous Vehicles

Head-to-Head autonomous racing at the limits of handling in the A2RL challenge

2026-02-09 · Simon Hoffmann, Simon Sagmeister, Tobias Betz, Joscha Bongard 외 arxiv

Autonomous racing presents a complex challenge involving multi-agent interactions between vehicles operating at the limit of performance and dynamics. As such, it provides a valuable research and testing environment for …

Autonomous Driving

Pushing The Limits of the Wiener Filter in Image Denoising

2023-03-27 · Clément Bled, François Pitié

As modern image denoiser networks have grown in size, their reported performance in popular real noise benchmarks such as DND and SIDD have now long outperformed classic non-deep learning denoisers such as Wiener and Wav…

DenoisingImage Denoising

US AISI and UK AISI Joint Pre-Deployment Test: Anthropic’s Claude 3.5 Sonnet (October 2024 Release)

2024-11-19 · NIST 2024 11 · US AI Safety Institute, UK AI Safety Institute

This technical report details a pre-deployment evaluation of Anthropic’s upgraded version of Claude 3.5 Sonnet, released October 22, 2024 (hereafter referred to as Sonnet 3.5 (new)). This evaluation was conducted jointly…