paper-with-me

홈 › Papers

Stabilising Explainability Fragility in Cybersecurity AI: The Impact and Mitigation of Multicollinearity in Public Benchmark Datasets

2026-05-21 · Ioannis J. Vourganas, Anna Lito Michala arxiv

This paper investigates a unexplored yet impactful vulnerability in AI explainability used in intrusion detection (IDS): multicollinearity-induced instability. Despite extensive reliance on post-hoc explainability tools such as SHAP or LIME, the impact of correlated features on explanation robustness is not evaluated. We introduce a formal theorem stating that multicollinearity inflates attribution variance. This demonstrates that explanations and feature importances are non-identifiable under multicollinearity. A suite of comprehensive experiments validates the theorem on a representative benchmark dataset, UNSW-NB15. Four widely used families of models are evaluated, including linear, tree-based, kernel, and neural, across full and pruned feature sets based on VIF and correlation thresholding. We propose the novel metric of Explanability Fragility Score and two novel methods to mitigate it with variable integration complexity. CAA-Filtering focuses on stabilising explanations by grouping attributions of trained models. SHARP is a novel training-time regularisation framework that penalises attribution instability, enabling controllable and monotonic improvement of explainability stability. The findings support stable predictive performance, using Kendall's τ to quantify instability across bootstrapped explanations. This work has direct implications for the trustworthiness and reproducibility of XAI in security-critical contexts, and motivates incorporating multicollinearity mitigations into the IDS pipelines, providing a set of guidelines for practitioners.

📄 PDF Abstract BibTeX arXiv:2605.22529

Code (0)

등록된 구현이 없습니다.

Tasks

Intrusion Detection

Similar Papers 제목 키워드 기반

Beyond Gradient-Based Attacks: Adversarial Robustness and Explainability Stability in Cybersecurity Classifiers

2026-07-02 · Mona Rajhans, Vishal Khawarey arxiv

Adversarial attacks on cybersecurity classifiers pose a dual threat: degrading predictions and destabilising the SHAP-based explanations that security analysts rely on to understand and triage alerts. We extend our prior…

Adversarial Robustness

Frontier AI's Impact on the Cybersecurity Landscape

2025-04-07 · Wenbo Guo, Yujin Potter, Tianneng Shi, Zhun Wang 외

As frontier AI advances rapidly, understanding its impact on cybersecurity and inherent risks is essential to ensuring safe AI evolution (e.g., guiding risk mitigation and informing policymakers). While some studies revi…

Automated CVE Analysis for Threat Prioritization and Impact Prediction

2023-09-06 · Ehsan Aghaei, Ehab Al-Shaer, Waseem Shadid, Xi Niu

The Common Vulnerabilities and Exposures (CVE) are pivotal information for proactive cybersecurity measures, including service patching, security hardening, and more. However, CVEs typically offer low-level, product-orie…

Prediction

A Novel Stochastic Epidemic Model with Application to COVID-19

2021-02-16 · Edilson F. Arruda, Rodrigo e Alvim Alexandre, Marcelo D. Fragoso, João B. R. do val 외

In this paper we propose a novel SEIR stochastic epidemic model. A distinguishing feature of this new model is that it allows us to consider a set up under general latency and infectious period distributions. To some ext…

Strategic commitments shape collective cybersecurity under AI inequality

2026-05-10 · Adeela Bashir, Zia Ush Shamszaman, Zhao Song, Matjaz Perc 외 arxiv

The growing integration of AI into cybersecurity is reshaping the balance between attackers and defenders. When access to advanced AI-enabled defence tools is uneven, resource-limited defenders may be unable to adopt eff…