Adversarial Robustness of Neural-Statistical Features in Detection of Generative Transformers
The detection of computer-generated text is an area of rapidly increasing significance as nascent generative models allow for efficient creation of compelling human-like text, which may be abused for the purposes of spam, disinformation, phishing, or online influence campaigns. Past work has studied detection of current state-of-the-art models, but despite a developing threat landscape, there has been minimal analysis of the robustness of detection methods to adversarial attacks. To this end, we evaluate neural and non-neural approaches on their ability to detect computer-generated text, their robustness against text adversarial attacks, and the impact that successful adversarial attacks have on human judgement of text quality. We find that while statistical features underperform neural features, statistical features provide additional adversarial robustness that can be leveraged in ensemble detection models. In the process, we find that previously effective complex phrasal features for detection of computer-generated text hold little predictive power against contemporary generative models, and identify promising statistical features to use instead. Finally, we pioneer the usage of $\Delta$MAUVE as a proxy measure for human judgement of adversarial text quality.
Code (1)
Tasks
Adversarial RobustnessAdversarial TextSimilar Papers 제목 키워드 기반
IDSGAN: Generative Adversarial Networks for Attack Generation against Intrusion Detection
As an essential tool in security, the intrusion detection system bears the responsibility of the defense to network attacks performed by malicious traffic. Nowadays, with the help of machine learning algorithms, intrusio…
Adversarial AttackIntrusion DetectionAdversarial Framework with Certified Robustness for Time-Series Domain via Statistical Features
Time-series data arises in many real-world applications (e.g., mobile health) and deep neural networks (DNNs) have shown great success in solving them. Despite their success, little is known about their robustness to adv…
Time SeriesTime Series AnalysisEvent Detection in Micro-PMU Data: A Generative Adversarial Network Scoring Method
A new data-driven method is proposed to detect events in the data streams from distribution-level phasor measurement units, a.k.a., micro-PMUs. The proposed method is developed by constructing unsupervised deep learning …
Anomaly DetectionEvent DetectionGenerative Adversarial NetworkCan Shape Structure Features Improve Model Robustness Under Diverse Adversarial Settings?
Recent studies show that convolutional neural networks (CNNs) are vulnerable under various settings, including adversarial attacks, common corruptions, and backdoor attacks. Motivated by the findings that human visua…
Edge DetectionGenerative Adversarial NetworkCausAdv: A Causal-based Framework for Detecting Adversarial Examples
Deep learning has led to tremendous success in many real-world applications of computer vision, thanks to sophisticated architectures such as Convolutional neural networks (CNNs). However, CNNs have been shown to be vuln…
Adversarial RobustnesscounterfactualCounterfactual Reasoning