paper-with-me

홈 › Papers

Research on Security Enhancement Methods for Adversarial Robust Large Language Model Intelligent Agents for Medical Decision-Making Tasks

2026-05-07 · Saisai Hu arxiv

Motivated by the challenge to improve the adversarial robustness, security, and trust of medical decision making intelligent agents, this study develops a full-link security enhancement framework, which describes "input risk perception - medical evidence constraint - knowledge consistency verification - decision confidence reweighting - security output control - adversarial feedback update." We propose ARSM-Agent and define a weighted joint objective consisting of decision accuracy loss, adversarial robustness loss, safety refusal loss, and knowledge consistency loss, with weights of 0.3, 0.3, 0.2, and 0.2, respectively. The whole medical decision formulation is implemented by multi-module collaborative linkage. We verify that the algorithm is more efficient than four baselines, including LLM-Agent, Retrieval-Agent, Filter-Agent, and Adv-Train-Agent. Under semantic perturbation, prompt injection, drug-name confusion, and false-evidence attacks, ARSM-Agent reduces the overall attack success rate to 8.7% and achieves a knowledge consistency score of 0.91. Ablation experiments quantify each module's contribution: removing risk perception, evidence retrieval, consistency verification, and confidence reweighting reduces accuracy by 6.7%, 9.1%, 7.6%, and 4.4%, respectively, and increases attack success rate by 13.8%, 11.1%, 8.6%, and 6.9%. The proposed approach addresses key security issues of medical decision making intelligent agents, obtains secure decision making in challenging scenarios, and provides reliable intelligent support for medical decision-making intelligent agents.

📄 PDF Abstract BibTeX arXiv:2605.08257

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial RobustnessDecision Making

Similar Papers 제목 키워드 기반

HASP: A High-Performance Adaptive Mobile Security Enhancement Against Malicious Speech Recognition

2018-09-04 · Zirui Xu, Fuxun Yu, ChenChen Liu, Xiang Chen

Nowadays, machine learning based Automatic Speech Recognition (ASR) technique has widely spread in smartphones, home devices, and public facilities. As convenient as this technology can be, a considerable security issue …

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Mobile Securityspeech-recognition+1

Adversarial Machine Learning Attacks and Defense Methods in the Cyber Security Domain

2020-07-05 · Ihai Rosenberg, Asaf Shabtai, Yuval Elovici, Lior Rokach

In recent years machine learning algorithms, and more specifically deep learning algorithms, have been widely used in many fields, including cyber security. However, machine learning systems are vulnerable to adversarial…

Adversarial AttackBIG-bench Machine Learning

Adversarial Security Attacks and Perturbations on Machine Learning and Deep Learning Methods

2019-07-17 · Arif Siddiqi

The ever-growing big data and emerging artificial intelligence (AI) demand the use of machine learning (ML) and deep learning (DL) methods. Cybersecurity also benefits from ML and DL methods for various types of applicat…

BIG-bench Machine Learning

Investigating cybersecurity incidents using large language models in latest-generation wireless networks

2025-04-14 · Leonid Legashev, Arthur Zhigalov

The purpose of research: Detection of cybersecurity incidents and analysis of decision support and assessment of the effectiveness of measures to counter information security threats based on modern generative models. Th…

Binary ClassificationData PoisoningFeature ImportanceLarge Language Model

Why Should Adversarial Perturbations be Imperceptible? Rethink the Research Paradigm in Adversarial NLP

2022-10-19 · Yangyi Chen, Hongcheng Gao, Ganqu Cui, Fanchao Qi 외

Textual adversarial samples play important roles in multiple subfields of NLP research, including security, evaluation, explainability, and data augmentation. However, most work mixes all these roles, obscuring the probl…

Data Augmentation