paper-with-me

홈 › Papers

RAGShield: Detecting Numerical Claim Manipulation in Government RAG Systems

2026-04-01 · KrishnaSaiReddy Patil arxiv

Retrieval-Augmented Generation (RAG) systems are deployed across federal agencies for citizen-facing tax guidance, benefits eligibility, and legal information, where a single incorrect number causes direct financial harm. This paper proves that all embedding-based RAG defenses share a fundamental blind spot: changing a tax deduction by $50,000 produces cosine similarity 0.9998, invisible to every known detection threshold. Across 174 manipulation pairs and two embedding models, the mean sensitivity gap is 1,459x. The blind spot is confirmed on real IRS documents.The root cause is that embeddings encode topic, not numerical precision. RAGShield sidesteps this by operating on extracted values directly: a pattern-based engine identifies dollar amounts and percentages in government text, links each value to its governing entity through two-pass context propagation (99.8% entity detection on 2,742 real IRS passages), and verifies every claim against a cross-source registry built from the corpus itself. A temporal tracker flags value changes that fall outside known government update schedules. On 430 attacks generated from real IRS document content, RAGShield detects every one (0.0% ASR, 95% CI [0%, 1%]) while embedding-based defenses miss 79-90% of the same attacks.

📄 PDF Abstract BibTeX arXiv:2604.00387

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Detecting False Claims in Low-Resource Regions: A Case Study of Caribbean Islands

2022-05-01 · CONSTRAINT (ACL) 2022 5 · Jason Lucas, Limeng Cui, Thai Le, Dongwon Lee

The COVID-19 pandemic has created threats to global health control. Misinformation circulated on social media and news outlets has undermined public trust towards Government and health agencies. This problem is further e…

Fact CheckingMisinformation

CLAIM: An Intent-Driven Multi-Agent Framework for Analyzing Manipulation in Courtroom Dialogues

2025-06-04 · Disha Sheshanarayana, Tanishka Magar, Ayushi Mittal, Neelam Chaplot

Courtrooms are places where lives are determined and fates are sealed, yet they are not impervious to manipulation. Strategic use of manipulation in legal jargon can sway the opinions of judges and affect the decisions. …

Decision MakingFairness

Prospects for inconsistency detection using large language models and sheaves

2024-01-30 · Steve Huntsman, Michael Robinson, Ludmilla Huntsman

We demonstrate that large language models can produce reasonable numerical ratings of the logical consistency of claims. We also outline a mathematical approach based on sheaf theory for lifting such ratings to hypertext…

Jurisprudence

Challenges and Opportunities in Information Manipulation Detection: An Examination of Wartime Russian Media

2022-05-24 · Chan Young Park, Julia Mendelsohn, Anjalie Field, Yulia Tsvetkov

NLP research on public opinion manipulation campaigns has primarily focused on detecting overt strategies such as fake news and disinformation. However, information manipulation in the ongoing Russia-Ukraine war exemplif…

Grand Challenge On Detecting Cheapfakes

2023-04-03 · Duc-Tien Dang-Nguyen, Sohail Ahmed Khan, Cise Midoglu, Michael Riegler 외

Cheapfake is a recently coined term that encompasses non-AI ("cheap") manipulations of multimedia content. Cheapfakes are known to be more prevalent than deepfakes. Cheapfake media can be created using editing software f…

Image Captioning