paper-with-me

홈 › Papers

Enhancing AI Interpretability and Safety through Localised Architectures

2026-06-06 · Ian Seet, Jonas Bozenhard, Simon Ostermann arxiv

Recent advances in generative AI, especially powerful Large Language Models (LLMs) and Large Reasoning Models (LRMs), raise concerns over the interpretability, safety and sustainability of these large and opaque AI models. The power of such architectures is derived not only from the scalability of deep neural networks, but also massively parallel hardware such as GPU clusters. The diffuse nature of deep neural networks gives them great function-approximation capability when provided with sufficient training data but imposes a cost in interpretability and computational efficiency. Observing that localised machine learning (ML) models tend to be more interpretable and computationally efficient than deep neural networks on small datasets, we reason by analogy that similar advantages may apply to specific localised hardware ML architectures. We argue that localised architectures with lower bandwidth but higher expressivity per node have the potential to be fundamentally more interpretable than deep neural networks running on GPU clusters while remaining competitive for smaller datasets. We then evaluate the suitability of various hardware ML paradigms for implementing such localised architectures and evaluate their per-node expressivity, energy efficiency and practical maturity of the technology required.

📄 PDF Abstract BibTeX arXiv:2606.07998

Code (0)

등록된 구현이 없습니다.

Tasks

Computational Efficiency

Similar Papers 제목 키워드 기반

Why Not Act on What You Know? Unleashing Safety Potential of LLMs via Self-Aware Guard Enhancement

2025-05-17 · Peng Ding, Jun Kuang, ZongYu Wang, Xuezhi Cao 외

Large Language Models (LLMs) have shown impressive capabilities across various tasks but remain vulnerable to meticulously crafted jailbreak attacks. In this paper, we identify a critical safety gap: while LLMs are adept…

Metrics that matter: Evaluating image quality metrics for medical image generation

2025-05-12 · Yash Deo, Yan Jia, Toni Lassila, William A. P. Smith 외

Evaluating generative models for synthetic medical imaging is crucial yet challenging, especially given the high standards of fidelity, anatomical accuracy, and safety required for clinical applications. Standard evaluat…

Image GenerationMedical Image Generation

Advancing Welding Defect Detection in Maritime Operations via Adapt-WeldNet and Defect Detection Interpretability Analysis

2025-08-01 · Kamal Basha S, Athira Nambiar arxiv

Weld defect detection is crucial for ensuring the safety and reliability of piping systems in the oil and gas industry, especially in challenging marine and offshore environments. Traditional non-destructive testing (NDT…

Transfer Learning

Going PLACES: Participatory Localized Red Teaming for Text-to-Image Safety in the Global South

2026-05-18 · Charvi Rastogi, Mukul Bhutani, Minsuk Kahng, Shamsuddeen Hassan Muhammad 외 arxiv

Despite the global deployment of text-to-image (T2I) models, their safety frameworks are largely calibrated to a Western-centric default, creating significant vulnerabilities for the rest of the world. To embrace cultura…

Red Teaming

VisionAD, a software package of performant anomaly detection algorithms, and Proportion Localised, an interpretable metric

2024-06-07 · Transactions of Machine Learning Research 2024 6 · Alexander D. J. Taylor, Phillip Tregidgo, Jonathan James Morrison, Neill D. F. Campbell

We release VisionAD, an anomaly detection library in the domain of images. The library forms the largest and most performant collection of such algorithms to date. Each algorithm is written through a standardised API, fo…

Anomaly DetectionBenchmarking