paper-with-me

홈 › Papers

RADAR: Robust AI-Text Detection via Adversarial Learning

2023-07-07 · NeurIPS 2023 11

Recent advances in large language models (LLMs) and the intensifying popularity of ChatGPT-like applications have blurred the boundary of high-quality text generation between humans and machines. However, in addition to the anticipated revolutionary changes to our technology and society, the difficulty of distinguishing LLM-generated texts (AI-text) from human-generated texts poses new challenges of misuse and fairness, such as fake content generation, plagiarism, and false accusations of innocent writers. While existing works show that current AI-text detectors are not robust to LLM-based paraphrasing, this paper aims to bridge this gap by proposing a new framework called RADAR, which jointly trains a robust AI-text detector via adversarial learning. RADAR is based on adversarial training of a paraphraser and a detector. The paraphraser's goal is to generate realistic content to evade AI-text detection. RADAR uses the feedback from the detector to update the paraphraser, and vice versa. Evaluated with 8 different LLMs (Pythia, Dolly 2.0, Palmyra, Camel, GPT-J, Dolly 1.0, LLaMA, and Vicuna) across 4 datasets, experimental results show that RADAR significantly outperforms existing AI-text detection methods, especially when paraphrasing is in place. We also identify the strong transferability of RADAR from instruction-tuned LLMs to other LLMs, and evaluate the improved capability of RADAR via GPT-3.5-Turbo.

📄 PDF Abstract BibTeX arXiv:2307.03838

Code (0)

등록된 구현이 없습니다.

Tasks

FairnessText DetectionText Generation

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
{Dispute@FaQ-s}How to file a dispute with Expedia? How to file a dispute with Expedia? To file a complaint against Expedia, first try contacting their customer service directly. You can reach them by phone at…
Multi-Head Attention 설명 없음
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Adam 설명 없음
Cosine Annealing Cosine Annealing is a type of learning rate schedule that has the effect of starting with a large learning rate that is relatively rapidly decreased to a minimum value before…
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

Constrained Contextual Bandit Learning for Adaptive Radar Waveform Selection

2021-03-09 · Charles E. Thornton, R. Michael Buehrer, Anthony F. Martone

A sequential decision process in which an adaptive radar system repeatedly interacts with a finite-state target channel is studied. The radar is capable of passively sensing the spectrum at regular intervals, which provi…

Thompson Sampling

Effective and Efficient Adversarial Detection for Vision-Language Models via A Single Vector

2024-10-30 · Youcheng Huang, Fengbin Zhu, Jingkun Tang, Pan Zhou 외

Visual Language Models (VLMs) are vulnerable to adversarial attacks, especially those from adversarial images, which is however under-explored in literature. To facilitate research on this critical safety problem, we fir…

RADAR: Retrieval-Augmented Detector with Adversarial Refinement for Robust Fake News Detection

2026-01-07 · Song-Duo Ma, Yi-Hung Liu, Hsin-Yu Lin, Pin-Yu Chen 외 arxiv

To efficiently combat the spread of LLM-generated misinformation, we present RADAR, a Retrieval-Augmented Detector with Adversarial Refinement for robust fake news detection. Our approach employs a generator that rewrite…

Fake News DetectionPassage Retrieval

Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text

2025-06-08 · Yize Cheng, Vinu Sankar Sadasivan, Mehrdad Saberi, Shoumik Saha 외

The increasing capabilities of Large Language Models (LLMs) have raised concerns about their misuse in AI-generated plagiarism and social engineering. While various AI-generated text detectors have been proposed to mitig…

Instruction Following

Generative Adversarial Network for Radar Signal Generation

2020-08-07 · Thomas Truong, Svetlana Yanushkevich

A major obstacle in radar based methods for concealed object detection on humans and seamless integration into security and access control system is the difficulty in collecting high quality radar signal data. Generative…

Generative Adversarial NetworkObjectobject-detectionObject Detection