paper-with-me

홈 › Papers

KGPA: Robustness Evaluation for Large Language Models via Cross-Domain Knowledge Graphs

2024-06-16 · Aihua Pei, Zehua Yang, Shunan Zhu, Ruoxi Cheng, Ju Jia, Lina Wang

Existing frameworks for assessing robustness of large language models (LLMs) overly depend on specific benchmarks, increasing costs and failing to evaluate performance of LLMs in professional domains due to dataset limitations. This paper proposes a framework that systematically evaluates the robustness of LLMs under adversarial attack scenarios by leveraging knowledge graphs (KGs). Our framework generates original prompts from the triplets of knowledge graphs and creates adversarial prompts by poisoning, assessing the robustness of LLMs through the results of these adversarial attacks. We systematically evaluate the effectiveness of this framework and its modules. Experiments show that adversarial robustness of the ChatGPT family ranks as GPT-4-turbo > GPT-4o > GPT-3.5-turbo, and the robustness of large language models is influenced by the professional domains in which they operate.

📄 PDF Abstract BibTeX arXiv:2406.10802

Code (1)

aika-wsd/KGPA 공식 구현

Tasks

Adversarial AttackAdversarial RobustnessKnowledge Graphs

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
15 Ways to Contact How can i speak to someone at Delta Airlines 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Residual Connection 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
BPE Byte Pair Encoding, or BPE, is a subword segmentation algorithm that encodes rare and unknown words as sequences of subword units. The intuition is that various word…
Adam 설명 없음

Similar Papers 제목 키워드 기반

Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input

2025-10-19 · Chenxu Li, Zhicai Wang, Yuan Sheng, Xingyu Zhu 외 arxiv

Multimodal Large Language Models (MLLMs) increasingly support dynamic image resolutions. However, current evaluation paradigms primarily assess semantic performance, overlooking the critical question of resolution robust…

DARE: Diverse Visual Question Answering with Robustness Evaluation

2024-09-26 · Hannah Sterz, Jonas Pfeiffer, Ivan Vulić

Vision Language Models (VLMs) extend remarkable capabilities of text-only large language models and vision-only models, and are able to learn from and process multi-modal vision-text input. While modern VLMs perform well…

image-classificationImage ClassificationImage-text matchingMultiple-choice+5

Multi-lingual Functional Evaluation for Large Language Models

2025-06-25 · Victor Ojewale, Inioluwa Deborah Raji, Suresh Venkatasubramanian

Multi-lingual competence in large language models is often evaluated via static data benchmarks such as Belebele, M-MMLU and M-GSM. However, these evaluations often fail to provide an adequate understanding of the practi…

BelebeleInstruction FollowingMathMMLU

Adversarial Moral Stress Testing of Large Language Models

2026-04-01 · Saeid Jamshidi, Foutse Khomh, Arghavan Moradi Dakhel, Amin Nikanjam 외 arxiv

Evaluating the ethical robustness of large language models (LLMs) deployed in software systems remains challenging, particularly under sustained adversarial user interaction. Existing safety benchmarks typically rely on …

On the Robustness of Knowledge Editing for Detoxification

2026-02-11 · Ming Dong, Shiyi Tang, Ziyan Peng, Guanyi Chen 외 arxiv

Knowledge-Editing-based (KE-based) detoxification has emerged as a promising approach for mitigating harmful behaviours in Large Language Models. Existing evaluations, however, largely rely on automatic toxicity classifi…

knowledge editing