paper-with-me

홈 › Papers

Towards Reliable Detection of LLM-Generated Texts: A Comprehensive Evaluation Framework with CUDRT

2024-06-13 · Zhen Tao, Yanfang Chen, Dinghao Xi, Zhiyu Li, Wei Xu

The increasing prevalence of large language models (LLMs) has significantly advanced text generation, but the human-like quality of LLM outputs presents major challenges in reliably distinguishing between human-authored and LLM-generated texts. Existing detection benchmarks are constrained by their reliance on static datasets, scenario-specific tasks (e.g., question answering and text refinement), and a primary focus on English, overlooking the diverse linguistic and operational subtleties of LLMs. To address these gaps, we propose CUDRT, a comprehensive evaluation framework and bilingual benchmark in Chinese and English, categorizing LLM activities into five key operations: Create, Update, Delete, Rewrite, and Translate. CUDRT provides extensive datasets tailored to each operation, featuring outputs from state-of-the-art LLMs to assess the reliability of LLM-generated text detectors. This framework supports scalable, reproducible experiments and enables in-depth analysis of how operational diversity, multilingual training sets, and LLM architectures influence detection performance. Our extensive experiments demonstrate the framework's capacity to optimize detection systems, providing critical insights to enhance reliability, cross-linguistic adaptability, and detection accuracy. By advancing robust methodologies for identifying LLM-generated texts, this work contributes to the development of intelligent systems capable of meeting real-world multilingual detection challenges. Source code and dataset are available at GitHub.

📄 PDF Abstract BibTeX arXiv:2406.09056

Code (2)

TaoZhen1110/CUDRT_Benchmark 공식 구현 pytorch
taozhen1110/cudrt 공식 구현 pytorch

Tasks

BenchmarkingLLM-generated Text DetectionQuestion AnsweringText DetectionText Generation

Methods 이 논문이 사용한 방법론

Focus 설명 없음

Similar Papers 제목 키워드 기반

The Science of Detecting LLM-Generated Texts

2023-02-04 · Ruixiang Tang, Yu-Neng Chuang, Xia Hu

The emergence of large language models (LLMs) has resulted in the production of LLM-generated texts that is highly sophisticated and almost indistinguishable from texts written by humans. However, this has also sparked c…

LLM-generated Text DetectionMisinformationText DetectionText Generation

Can AI-Generated Persuasion Be Detected? Persuaficial Benchmark and AI vs. Human Linguistic Differences

2026-01-08 · Arkadiusz Modzelewski, Paweł Golik, Anna Kołos, Giovanni Da San Martino arxiv

Large Language Models (LLMs) can generate highly persuasive text, raising concerns about their misuse for propaganda, manipulation, and other harmful purposes. This leads us to our central question: Is LLM-generated pers…

Intrinsic Task-based Evaluation for Referring Expression Generation

2024-02-12 · Guanyi Chen, Fahime Same, Kees Van Deemter

Recently, a human evaluation study of Referring Expression Generation (REG) models had an unexpected conclusion: on \textsc{webnlg}, Referring Expressions (REs) generated by the state-of-the-art neural models were not on…

Referring ExpressionReferring expression generationText Generation

RKadiyala at SemEval-2024 Task 8: Black-Box Word-Level Text Boundary Detection in Partially Machine Generated Texts

2024-10-22 · Ram Mohan Rao Kadiyala

With increasing usage of generative models for text generation and widespread use of machine generated texts in various domains, being able to distinguish between human written and machine generated texts is a significan…

Boundary DetectionSentenceText Generation

GPTZero: Robust Detection of LLM-Generated Texts

2026-02-13 · George Alexandru Adam, Alexander Cui, Edwin Thomas, Emily Napier 외 arxiv

While historical considerations surrounding text authenticity revolved primarily around plagiarism, the advent of large language models (LLMs) has introduced a new challenge: distinguishing human-authored from AI-generat…

Red Teaming