paper-with-me

홈 › Papers

Disappearing Ink: Obfuscation Breaks N-gram Code Watermarks in Theory and Practice

2025-07-07 · Gehao Zhang, Eugene Bagdasarian, Juan Zhai, Shiqing Ma

Distinguishing AI-generated code from human-written code is becoming crucial for tasks such as authorship attribution, content tracking, and misuse detection. Based on this, N-gram-based watermarking schemes have emerged as prominent, which inject secret watermarks to be detected during the generation. However, their robustness in code content remains insufficiently evaluated. Most claims rely solely on defenses against simple code transformations or code optimizations as a simulation of attack, creating a questionable sense of robustness. In contrast, more sophisticated schemes already exist in the software engineering world, e.g., code obfuscation, which significantly alters code while preserving functionality. Although obfuscation is commonly used to protect intellectual property or evade software scanners, the robustness of code watermarking techniques against such transformations remains largely unexplored. In this work, we formally model the code obfuscation and prove the impossibility of N-gram-based watermarking's robustness with only one intuitive and experimentally verified assumption, distribution consistency, satisfied. Given the original false positive rate of the watermarking detection, the ratio that the detector failed on the watermarked code after obfuscation will increase to 1 - fpr. The experiments have been performed on three SOTA watermarking schemes, two LLMs, two programming languages, four code benchmarks, and four obfuscators. Among them, all watermarking detectors show coin-flipping detection abilities on obfuscated codes (AUROC tightly surrounds 0.5). Among all models, watermarking schemes, and datasets, both programming languages own obfuscators that can achieve attack effects with no detection AUROC higher than 0.6 after the attack. Based on the theoretical and practical observations, we also proposed a potential path of robust code watermarking.

📄 PDF Abstract BibTeX arXiv:2507.05512

Code (0)

등록된 구현이 없습니다.

Tasks

Authorship Attribution

Similar Papers 제목 키워드 기반

Plentiful Jailbreaks with String Compositions

2024-11-01 · Brian R. Y. Huang

Large language models (LLMs) remain vulnerable to a slew of adversarial attacks and jailbreaking methods. One common approach employed by white-hat attackers, or red-teamers, is to process model inputs and outputs using …

Rethinking White-Box Watermarks on Deep Learning Models under Neural Structural Obfuscation

2023-03-17 · Yifan Yan, Xudong Pan, Mi Zhang, Min Yang

Copyright protection for deep neural networks (DNNs) is an urgent need for AI corporations. To trace illegally distributed model copies, DNN watermarking is an emerging technique for embedding and verifying secret identi…

A Systematic Study of Code Obfuscation Against LLM-based Vulnerability Detection

2025-12-18 · Xiao Li, Yue Li, Hao Wu, Yue Zhang 외 arxiv

As large language models (LLMs) are increasingly adopted for code vulnerability detection, their reliability and robustness across diverse vulnerability types have become a pressing concern. In traditional adversarial se…

Vulnerability Detection

Analyzing Chain of Thought (CoT) Approaches in Control Flow Code Deobfuscation Tasks

2026-04-16 · Seyedreza Mohseni, Sarvesh Baskar, Edward Raff, Manas Gaur arxiv

Code deobfuscation is the task of recovering a readable version of a program while preserving its original behavior. In practice, this often requires days or even months of manual work with complex and expensive analysis…

The Code Barrier: What LLMs Actually Understand?

2025-04-14 · Serge Lionel Nikiema, Jordan Samhi, Abdoul Kader Kaboré, Jacques Klein 외

Understanding code represents a core ability needed for automating software development tasks. While foundation models like LLMs show impressive results across many software engineering challenges, the extent of their tr…