paper-with-me

홈 › Papers

Modeling the Attack: Detecting AI-Generated Text by Quantifying Adversarial Perturbations

2025-09-22 · Lekkala Sai Teja, Annepaka Yadagiri, Sangam Sai Anish, Siva Gopala Krishna Nuthakki, Partha Pakray arxiv

The growth of highly advanced Large Language Models (LLMs) constitutes a huge dual-use problem, making it necessary to create dependable AI-generated text detection systems. Modern detectors are notoriously vulnerable to adversarial attacks, with paraphrasing standing out as an effective evasion technique that foils statistical detection. This paper presents a comparative study of adversarial robustness, first by quantifying the limitations of standard adversarial training and then by introducing a novel, significantly more resilient detection framework: Perturbation-Invariant Feature Engineering (PIFE), a framework that enhances detection by first transforming input text into a standardized form using a multi-stage normalization pipeline, it then quantifies the transformation's magnitude using metrics like Levenshtein distance and semantic similarity, feeding these signals directly to the classifier. We evaluate both a conventionally hardened Transformer and our PIFE-augmented model against a hierarchical taxonomy of character-, word-, and sentence-level attacks. Our findings first confirm that conventional adversarial training, while resilient to syntactic noise, fails against semantic attacks, an effect we term "semantic evasion threshold", where its True Positive Rate at a strict 1% False Positive Rate plummets to 48.8%. In stark contrast, our PIFE model, which explicitly engineers features from the discrepancy between a text and its canonical form, overcomes this limitation. It maintains a remarkable 82.6% TPR under the same conditions, effectively neutralizing the most sophisticated semantic attacks. This superior performance demonstrates that explicitly modeling perturbation artifacts, rather than merely training on them, is a more promising path toward achieving genuine robustness in the adversarial arms race.

📄 PDF Abstract BibTeX arXiv:2510.02319

Code (0)

등록된 구현이 없습니다.

Tasks

Adversarial RobustnessSemantic SimilarityFeature EngineeringText Detection

Similar Papers 제목 키워드 기반

Relative Hausdorff Distance for Network Analysis

2019-06-12 · Sinan G. Aksoy, Kathleen E. Nowak, Emilie Purvine, Stephen J. Young

Similarity measures are used extensively in machine learning and data science algorithms. The newly proposed graph Relative Hausdorff (RH) distance is a lightweight yet nuanced similarity measure for quantifying the clos…

DSIPA: Detecting LLM-Generated Texts via Sentiment-Invariant Patterns Divergence Analysis

2026-04-29 · Siyuan Li, Aodu Wulianghai, Guangyan Li, Xi Lin 외 arxiv

The rapid advancement of large language models (LLMs) presents new security challenges, particularly in detecting machine-generated text used for misinformation, impersonation, and content forgery. Most existing detectio…

Amplifying Training Data Exposure through Fine-Tuning with Pseudo-Labeled Memberships

2024-02-19 · Myung Gyo Oh, Hong Eun Ahn, Leo Hyun Park, Taekyoung Kwon

Neural language models (LMs) are vulnerable to training data extraction attacks due to data memorization. This paper introduces a novel attack scenario wherein an attacker adversarially fine-tunes pre-trained LMs to ampl…

Memorization

TCAB: A Large-Scale Text Classification Attack Benchmark

2022-10-21 · Kalyani Asthana, Zhouhang Xie, Wencong You, Adam Noack 외

We introduce the Text Classification Attack Benchmark (TCAB), a dataset for analyzing, understanding, detecting, and labeling adversarial attacks against text classifiers. TCAB includes 1.5 million attack instances, gene…

Abuse DetectionClassificationSentiment Analysistext-classification+1

ARGUS: Context-Based Detection of Stealthy IoT Infiltration Attacks

2023-02-15 · Phillip Rieger, Marco Chilese, Reham Mohamed, Markus Miettinen 외

IoT application domains, device diversity and connectivity are rapidly growing. IoT devices control various functions in smart homes and buildings, smart cities, and smart factories, making these devices an attractive ta…

Intrusion DetectionSelf-Learning