paper-with-me

홈 › Papers

Paraphrasing Attack Resilience of Various AI-Generated Text Detection Methods

2026-05-14 · Andrii Shportko, Inessa Verbitsky arxiv

The recent large-scale emergence of LLMs has left an open space for dealing with their consequences, such as plagiarism or the spread of false information on the Internet. Coupling this with the rise of AI detector bypassing tools, reliable machine-generated text detection is in increasingly high demand. We investigate the paraphrasing attack resilience of various machine-generated text detection methods, evaluating three approaches: fine-tuned RoBERTa, Binoculars, and text feature analysis, along with their ensembles using Random Forest classifiers. We discovered that Binoculars-inclusive ensembles yield the strongest results, but they also suffer the most significant losses during attacks. In this paper, we present the dichotomy of performance versus resilience in the world of AI text detection, which complicates the current perception of reliability among state-of-the-art techniques.

📄 PDF Abstract BibTeX arXiv:2605.14240

Code (0)

등록된 구현이 없습니다.

Tasks

Text Detection

Similar Papers 제목 키워드 기반

Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated Text

2025-06-08 · Yize Cheng, Vinu Sankar Sadasivan, Mehrdad Saberi, Shoumik Saha 외

The increasing capabilities of Large Language Models (LLMs) have raised concerns about their misuse in AI-generated plagiarism and social engineering. While various AI-generated text detectors have been proposed to mitig…

Instruction Following

Evaluating adversarial attacks against multiple fact verification systems

2019-11-01 · IJCNLP 2019 11 · James Thorne, Andreas Vlachos, Christos Christodoulopoulos, Arpit Mittal

Automated fact verification has been progressing owing to advancements in modeling and availability of large datasets. Due to the nature of the task, it is critical to understand the vulnerabilities of these systems agai…

Fact Verification

Revisiting the Robustness of Watermarking to Paraphrasing Attacks

2024-11-08 · Saksham Rastogi, Danish Pruthi

Amidst rising concerns about the internet being proliferated with content generated from language models (LMs), watermarking is seen as a principled way to certify whether text was generated from a model. Many recent wat…

OUTFOX: LLM-Generated Essay Detection Through In-Context Learning with Adversarially Generated Examples

2023-07-21 · Ryuto Koike, Masahiro Kaneko, Naoaki Okazaki

Large Language Models (LLMs) have achieved human-level fluency in text generation, making it difficult to distinguish between human-written and LLM-generated texts. This poses a growing risk of misuse of LLMs and demands…

Adversarial AttackAdversarial Attack DetectionDeepFake DetectionIn-Context Learning+4

Can AI-Generated Text be Reliably Detected?

2023-03-17 · Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang 외

Large Language Models (LLMs) perform impressively well in various applications. However, the potential for misuse of these models in activities such as plagiarism, generating fake news, and spamming has raised concern ab…

Language ModellingLarge Language ModelQuestion AnsweringText Generation