paper-with-me

홈 › Papers

AVISE: Framework for Evaluating the Security of AI Systems

2026-04-22 · Mikko Lempinen, Joni Kemppainen, Niklas Raesalmi arxiv

As artificial intelligence (AI) systems are increasingly deployed across critical domains, their security vulnerabilities pose growing risks of high-profile exploits and consequential system failures. Yet systematic approaches to evaluating AI security remain underdeveloped. In this paper, we introduce AVISE (AI Vulnerability Identification and Security Evaluation), a modular open-source framework for identifying vulnerabilities in and evaluating the security of AI systems and models. As a demonstration of the framework, we extend the theory-of-mind-based multi-turn Red Queen attack into an Adversarial Language Model (ALM) augmented attack and develop an automated Security Evaluation Test (SET) for discovering jailbreak vulnerabilities in language models. The SET comprises 25 test cases and an Evaluation Language Model (ELM) that determines whether each test case was able to jailbreak the target model, achieving 92% accuracy, an F1-score of 0.91, and a Matthews correlation coefficient of 0.83. We evaluate nine recently released language models of diverse sizes with the SET and find that all are vulnerable to the augmented Red Queen attack to varying degrees. AVISE provides researchers and industry practitioners with an extensible foundation for developing and deploying automated SETs, offering a concrete step toward more rigorous and reproducible AI security evaluation.

📄 PDF Abstract BibTeX arXiv:2604.20833

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

NaviSense: A Multimodal Assistive Mobile application for Object Retrieval by Persons with Visual Impairment

2025-09-23 · Ajay Narayanan Sridhar, Fuli Qiao, Nelson Daniel Troncoso Aldas, Yanpei Shi 외 arxiv

People with visual impairments often face significant challenges in locating and retrieving objects in their surroundings. Existing assistive technologies present a trade-off: systems that offer precise guidance typicall…

Object RecognitionObject Detection

MAVISEG: Manifold Propagation and Visual Prototypes for Zero-Shot Open-Vocabulary Segmentation in Diffusion Transformers

2026-08-06 · Rajatsubhra Chakraborty, Xujun Che, Ritabrata Chakraborty, Xi Niu 외 arxiv

Text-to-image diffusion transformers learn about objects and scenes by learning to generate them, making them strong candidates for training-free zero-shot open-vocabulary semantic segmentation. State-of-the-art attribut…

Semantic Segmentation

Evaluating the Cybersecurity Risk of Real World, Machine Learning Production Systems

2021-07-05 · Ron Bitton, Nadav Maman, Inderjeet Singh, Satoru Momiyama 외

Although cyberattacks on machine learning (ML) production systems can be harmful, today, security practitioners are ill equipped, lacking methodologies and tactical tools that would allow them to analyze the security ris…

BIG-bench Machine LearningGraph Generation

Audio-Visual Instance Segmentation

2023-10-28 · CVPR 2025 1 · Ruohao Guo, Xianghua Ying, Yaru Chen, Dantong Niu 외

In this paper, we propose a new multi-modal task, termed audio-visual instance segmentation (AVIS), which aims to simultaneously identify, segment and track individual sounding object instances in audible videos. To faci…

Instance SegmentationSegmentationSemantic SegmentationSound Source Localization

Cyber Security Requirements for Platforms Enhancing AI Reproducibility

2023-09-27 · Polra Victor Falade

Scientific research is increasingly reliant on computational methods, posing challenges for ensuring research reproducibility. This study focuses on the field of artificial intelligence (AI) and introduces a new framewor…