paper-with-me

홈 › Papers

GPT-who: An Information Density-based Machine-Generated Text Detector

2023-10-09 · Saranya Venkatraman, Adaku Uchendu, Dongwon Lee

The Uniform Information Density (UID) principle posits that humans prefer to spread information evenly during language production. We examine if this UID principle can help capture differences between Large Language Models (LLMs)-generated and human-generated texts. We propose GPT-who, the first psycholinguistically-inspired domain-agnostic statistical detector. This detector employs UID-based features to model the unique statistical signature of each LLM and human author for accurate detection. We evaluate our method using 4 large-scale benchmark datasets and find that GPT-who outperforms state-of-the-art detectors (both statistical- & non-statistical) such as GLTR, GPTZero, DetectGPT, OpenAI detector, and ZeroGPT by over $20$% across domains. In addition to better performance, it is computationally inexpensive and utilizes an interpretable representation of text articles. We find that GPT-who can distinguish texts generated by very sophisticated LLMs, even when the overlying text is indiscernible. UID-based measures for all datasets and code are available at https://github.com/saranya-venkatraman/gpt-who.

📄 PDF Abstract BibTeX arXiv:2310.06202

Code (1)

saranya-venkatraman/gpt-who 공식 구현 pytorch

Tasks

ArticlesAuthorship Attribution

Similar Papers 제목 키워드 기반

TextFake: Benchmarking AI-Generated Image Detection on Text-Rich Images

2026-05-31 · Yuning Zhang, Changtao Miao, Mingyu Liao, Tingyu Liu 외 arxiv

Recent AI-generated image (AIGI) detectors perform well on natural-image benchmarks, but their behavior on text-rich forgeries, such as fabricated screenshots, documents, and news pages prevalent in misinformation, remai…

Are AI Detectors Good Enough? A Survey on Quality of Datasets With Machine-Generated Texts

2024-10-18 · German Gritsai, Anastasia Voznyuk, Andrey Grabovoy, Yury Chekhovich

The rapid development of autoregressive Large Language Models (LLMs) has significantly improved the quality of generated texts, necessitating reliable machine-generated text detectors. A huge number of detectors and coll…

Smaller Language Models are Better Black-box Machine-Generated Text Detectors

2023-05-17 · Niloofar Mireshghallah, Justus Mattern, Sicun Gao, Reza Shokri 외

With the advent of fluent generative language models that can produce convincing utterances very similar to those written by humans, distinguishing whether a piece of text is machine-generated or human-written becomes mo…

Misinformation

Adapting Fake News Detection to the Era of Large Language Models

2023-11-02 · Jinyan Su, Claire Cardie, Preslav Nakov

In the age of large language models (LLMs) and the widespread adoption of AI-driven content creation, the landscape of information dissemination has witnessed a paradigm shift. With the proliferation of both human-writte…

ArticlesFake News Detection

On the Zero-Shot Generalization of Machine-Generated Text Detectors

2023-10-08 · Xiao Pu, Jingyu Zhang, Xiaochuang Han, Yulia Tsvetkov 외

The rampant proliferation of large language models, fluent enough to generate text indistinguishable from human-written language, gives unprecedented importance to the detection of machine-generated text. This work is mo…

Zero-shot Generalization