paper-with-me

홈 › Papers

LLM Robustness Against Misinformation in Biomedical Question Answering

2024-10-27 · Alexander Bondarenko, Adrian Viehweger

The retrieval-augmented generation (RAG) approach is used to reduce the confabulation of large language models (LLMs) for question answering by retrieving and providing additional context coming from external knowledge sources (e.g., by adding the context to the prompt). However, injecting incorrect information can mislead the LLM to generate an incorrect answer. In this paper, we evaluate the effectiveness and robustness of four LLMs against misinformation - Gemma 2, GPT-4o-mini, Llama~3.1, and Mixtral - in answering biomedical questions. We assess the answer accuracy on yes-no and free-form questions in three scenarios: vanilla LLM answers (no context is provided), "perfect" augmented generation (correct context is provided), and prompt-injection attacks (incorrect context is provided). Our results show that Llama 3.1 (70B parameters) achieves the highest accuracy in both vanilla (0.651) and "perfect" RAG (0.802) scenarios. However, the accuracy gap between the models almost disappears with "perfect" RAG, suggesting its potential to mitigate the LLM's size-related effectiveness differences. We further evaluate the ability of the LLMs to generate malicious context on one hand and the LLM's robustness against prompt-injection attacks on the other hand, using metrics such as attack success rate (ASR), accuracy under attack, and accuracy drop. As adversaries, we use the same four LLMs (Gemma 2, GPT-4o-mini, Llama 3.1, and Mixtral) to generate incorrect context that is injected in the target model's prompt. Interestingly, Llama is shown to be the most effective adversary, causing accuracy drops of up to 0.48 for vanilla answers and 0.63 for "perfect" RAG across target models. Our analysis reveals that robustness rankings vary depending on the evaluation measure, highlighting the complexity of assessing LLM resilience to adversarial attacks.

📄 PDF Abstract BibTeX arXiv:2410.21330

Code (1)

alebondarenko/llm-robustness 공식 구현

Tasks

MisinformationQuestion AnsweringRAGRetrieval-augmented Generation

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Residual Connection 설명 없음
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

Attacking Open-domain Question Answering by Injecting Misinformation

2021-10-15 · Liangming Pan, Wenhu Chen, Min-Yen Kan, William Yang Wang

With a rise in false, inaccurate, and misleading information in propaganda, news, and social media, real-world Question Answering (QA) systems face the challenges of synthesizing and reasoning over misinformation-pollute…

MisinformationOpen-Domain Question AnsweringQuestion Answering

Do Images Speak Louder than Words? Investigating the Effect of Textual Misinformation in VLMs

2026-01-27 · Chi Zhang, Wenxuan Ding, Jiale Liu, Mingrui Wu 외 arxiv

Vision-Language Models (VLMs) have shown strong multimodal reasoning capabilities on Visual-Question-Answering (VQA) benchmarks. However, their robustness against textual misinformation remains under-explored. While exis…

Multimodal Reasoning

Open-Domain Question-Answering for COVID-19 and Other Emergent Domains

2021-10-13 · EMNLP (ACL) 2021 11 · Sharon Levy, Kevin Mo, Wenhan Xiong, William Yang Wang

Since late 2019, COVID-19 has quickly emerged as the newest biomedical domain, resulting in a surge of new information. As with other emergent domains, the discussion surrounding the topic has been rapidly changing, lead…

DiversityMisinformationOpen-Domain Question AnsweringQuestion Answering+1

Combining LLMs and Knowledge Graphs to Reduce Hallucinations in Question Answering

2024-09-06 · Larissa Pusch, Tim O. F. Conrad

Advancements in natural language processing have revolutionized the way we can interact with digital information systems, such as databases, making them more accessible. However, challenges persist, especially when accur…

HallucinationKnowledge GraphsMisinformationNatural Language Queries+2

Hierarchical Representation-based Dynamic Reasoning Network for Biomedical Question Answering

2022-10-01 · COLING 2022 10 · Jianguo Mao, Jiyuan Zhang, Zengfeng Zeng, Weihua Peng 외

Recently, Biomedical Question Answering (BQA) has attracted growing attention due to its application value and technical challenges. Most existing works treat it as a semantic matching task that predicts answers by compu…

MedQAQuestion Answering