paper-with-me

홈 › Papers

CleanBase: Detecting Malicious Documents in RAG Knowledge Databases

2026-05-01 · Weifei Jin, Xilong Wang, Wei Zou, Jinyuan Jia, Neil Gong arxiv

Retrieval-augmented generation (RAG) is vulnerable to prompt injection attacks, in which an adversary inserts malicious documents containing carefully crafted injected prompts into the knowledge database. When a user issues a question targeted by the attack, the RAG system may retrieve these malicious documents, whose injected prompts mislead it into generating attacker-specified answers, thereby compromising the integrity of the RAG system. In this work, we propose CleanBase, a method to detect malicious documents within a knowledge database. Our key insight is that malicious documents crafted for the same attack-targeted questions often exhibit high semantic similarity, as attackers deliberately make them consistent to improve attack success rates. Accordingly, CleanBase constructs a similarity graph over the knowledge database, where each node represents a document and an edge connects two nodes if their semantic similarity--computed using an embedding model--exceeds a statistically determined threshold. Due to their inherent similarity, malicious documents tend to form cliques within this graph. CleanBase detects such cliques and flags the corresponding documents as malicious. We theoretically derive upper bounds on CleanBase's false positive and false negative rates and empirically validate its effectiveness. Experimental results across multiple datasets and prompt injection attacks demonstrate that CleanBase accurately detects malicious documents and effectively safeguards RAG systems. Our source code is available at https://github.com/WeifeiJin/CleanBase.

📄 PDF Abstract BibTeX arXiv:2605.00460

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic Similarity

Similar Papers 제목 키워드 기반

Constructing Large Proposition Databases

2012-05-01 · LREC 2012 5 · Peter Exner, Pierre Nugues

With the advent of massive online encyclopedic corpora such as Wikipedia, it has become possible to apply a systematic analysis to a wide range of documents covering a significant part of human knowledge. Using semantic …

Dependency ParsingMachine TranslationQuestion AnsweringSemantic Dependency Parsing+1

RAGSieve: Self-Referenced Local Contrast for Knowledge-Poison Detection in Retrieval-Augmented Generation

2026-08-13 · Xinlong Xu, Yoshua Y. Li arxiv

Retrieval-augmented generation treats an external corpus as inference evidence, allowing injected documents to promote attacker-chosen claims. Existing detectors depend on trusted references, specific attack artifacts, o…

EntSQL: A Benchmark for Grounding Text-to-SQL in Long-Context Enterprise Knowledge

2026-06-02 · Chengxi Liao, Tao Xu, Zulong Chen, Chuanfei Xu 외 arxiv

Text-to-SQL enables natural language access to databases, and recent LLMs have substantially advanced its capabilities. Existing benchmarks such as Spider, BIRD, and Spider~2.0 evaluate schema generalization, large-scale…

Detecting Social Media Manipulation in Low-Resource Languages

2020-11-10 · Samar Haider, Luca Luceri, Ashok Deb, Adam Badawy 외

Social media have been deliberately used for malicious purposes, including political manipulation and disinformation. Most research focuses on high-resource languages. However, malicious actors share content across count…

Transfer Learning

FantasyID: A dataset for detecting digital manipulations of ID-documents

2025-07-28 · Pavel Korshunov, Amir Mohammadi, Vidit Vidit, Christophe Ecabert 외 arxiv

Advancements in image generation led to the availability of easy-to-use tools for malicious actors to create forged images. These tools pose a serious threat to the widespread Know Your Customer (KYC) applications, requi…

Image Generation