paper-with-me

홈 › Papers

Manifold of Failure: Behavioral Attraction Basins in Language Models

2026-02-25 · Sarthak Munshi, Manish Bhatt, Vineeth Sai Narajala, Idan Habler, Ammar Al-Kahfah, Ken Huang, Blake Gatto arxiv

While prior work has focused on projecting adversarial examples back onto the manifold of natural data to restore safety, we argue that a comprehensive understanding of AI safety requires characterizing the unsafe regions themselves. This paper introduces a framework for systematically mapping the Manifold of Failure in Large Language Models (LLMs). We reframe the search for vulnerabilities as a quality diversity problem, using MAP-Elites to illuminate the continuous topology of these failure regions, which we term behavioral attraction basins. Our quality metric, Alignment Deviation, guides the search towards areas where the model's behavior diverges most from its intended alignment. Across three LLMs: Llama-3-8B, GPT-OSS-20B, and GPT-5-Mini, we show that MAP-Elites achieves up to 63% behavioral coverage, discovers up to 370 distinct vulnerability niches, and reveals dramatically different model-specific topological signatures: Llama-3-8B exhibits a near-universal vulnerability plateau (mean Alignment Deviation 0.93), GPT-OSS-20B shows a fragmented landscape with spatially concentrated basins (mean 0.73), and GPT-5-Mini demonstrates strong robustness with a ceiling at 0.50. Our approach produces interpretable, global maps of each model's safety landscape that no existing attack method (GCG, PAIR, or TAP) can provide, shifting the paradigm from finding discrete failures to understanding their underlying structure.

📄 PDF Abstract BibTeX arXiv:2602.22291

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Detecting Invariant Manifolds in ReLU-Based RNNs

2025-10-04 · Lukas Eisenmann, Alena Brändle, Zahra Monfared, Daniel Durstewitz arxiv

Recurrent Neural Networks (RNNs) have found widespread applications in machine learning for time series prediction and dynamical systems reconstruction, and experienced a recent renaissance with improved training algorit…

Time Series Prediction

Topological properties of basins of attraction and expressiveness of width bounded neural networks

2020-11-10 · Hans-Peter Beise, Steve Dias Da Cruz

In Radhakrishnan et al. [2020], the authors empirically show that autoencoders trained with usual SGD methods shape out basins of attraction around their training data. We consider network functions of width not exceedin…

Visualising Basins of Attraction for the Cross-Entropy and the Squared Error Neural Network Loss Functions

2019-01-08 · Anna Sergeevna Bosman, Andries Engelbrecht, Mardé Helbig

Quantification of the stationary points and the associated basins of attraction of neural network loss surfaces is an important step towards a better understanding of neural network loss surfaces at large. This work prop…

Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data

2026-04-29 · Bao Pham, Mohammed J. Zaki, Luca Ambrogioni, Dmitry Krotov 외 arxiv

When do language diffusion models memorize their training data, and how to quantitatively assess their true generative regime? We address these questions by showing that Uniform-based Discrete Diffusion Models (UDDMs) fu…

Deep Learning-based Analysis of Basins of Attraction

2023-09-27 · David Valle, Alexandre Wagemakers, Miguel A. F. Sanjuán

This research addresses the challenge of characterizing the complexity and unpredictability of basins within various dynamical systems. The main focus is on demonstrating the efficiency of convolutional neural networks (…

Deep Learning