paper-with-me

홈 › Papers

Improving Robustness Estimates in Natural Language Explainable AI though Synonymity Weighted Similarity Measures

2025-01-02 · Christopher Burger

Explainable AI (XAI) has seen a surge in recent interest with the proliferation of powerful but intractable black-box models. Moreover, XAI has come under fire for techniques that may not offer reliable explanations. As many of the methods in XAI are themselves models, adversarial examples have been prominent in the literature surrounding the effectiveness of XAI, with the objective of these examples being to alter the explanation while maintaining the output of the original model. For explanations in natural language, it is natural to use measures found in the domain of information retrieval for use with ranked lists to guide the adversarial XAI process. We show that the standard implementation of these measures are poorly suited for the comparison of explanations in adversarial XAI and amend them by using information that is discarded, the synonymity of perturbed words. This synonymity weighting produces more accurate estimates of the actual weakness of XAI methods to adversarial examples.

📄 PDF Abstract BibTeX arXiv:2501.01516

Code (0)

등록된 구현이 없습니다.

Tasks

Information Retrieval

Similar Papers 제목 키워드 기반

Adversarially Robust and Explainable Model Compression with On-Device Personalization for Text Classification

2021-01-10 · Yao Qiang, Supriya Tumkur Suresh Kumar, Marco Brocanelli, Dongxiao Zhu

On-device Deep Neural Networks (DNNs) have recently gained more attention due to the increasing computing power of the mobile devices and the number of applications in Computer Vision (CV), Natural Language Processing (N…

Adversarial RobustnessGeneral ClassificationKnowledge DistillationModel Compression+2

Towards Robust and Accurate Stability Estimation of Local Surrogate Models in Text-based Explainable AI

2025-01-03 · Christopher Burger, Charles Walter, Thai Le, Lingwei Chen

Recent work has investigated the concept of adversarial attacks on explainable AI (XAI) in the NLP domain with a focus on examining the vulnerability of local surrogate methods such as Lime to adversarial perturbations o…

Adversarial Robustness

TactEx: An Explainable Multimodal Robotic Interaction Framework for Human-Like Touch and Hardness Estimation

2026-02-21 · Felix Verstraete, Lan Wei, Wen Fan, Dandan Zhang arxiv

Accurate perception of object hardness is essential for safe and dexterous contact-rich robotic manipulation. Here, we present TactEx, an explainable multimodal robotic interaction framework that unifies vision, touch, a…

Attack and defense techniques in large language models: A survey and new perspectives

2025-05-02 · Zhiyu Liao, Kang Chen, Yuanguo Lin, Kangkang Li 외

Large Language Models (LLMs) have become central to numerous natural language processing tasks, but their vulnerabilities present significant security and ethical challenges. This systematic survey explores the evolving …

Survey

Explainable Natural Language Processing with Matrix Product States

2021-12-16 · Jirawat Tangpanitanon, Chanatip Mangkang, Pradeep Bhadola, Yuichiro Minato 외

Despite empirical successes of recurrent neural networks (RNNs) in natural language processing (NLP), theoretical understanding of RNNs is still limited due to intrinsically complex non-linear computations. We systematic…

Sentiment Analysis