paper-with-me

홈 › Papers

Is ChatGPT Involved in Texts? Measure the Polish Ratio to Detect ChatGPT-Generated Text

2023-07-21 · Lingyi Yang, Feng Jiang, Haizhou Li

The remarkable capabilities of large-scale language models, such as ChatGPT, in text generation have impressed readers and spurred researchers to devise detectors to mitigate potential risks, including misinformation, phishing, and academic dishonesty. Despite this, most previous studies have been predominantly geared towards creating detectors that differentiate between purely ChatGPT-generated texts and human-authored texts. This approach, however, fails to work on discerning texts generated through human-machine collaboration, such as ChatGPT-polished texts. Addressing this gap, we introduce a novel dataset termed HPPT (ChatGPT-polished academic abstracts), facilitating the construction of more robust detectors. It diverges from extant corpora by comprising pairs of human-written and ChatGPT-polished abstracts instead of purely ChatGPT-generated texts. Additionally, we propose the "Polish Ratio" method, an innovative measure of the degree of modification made by ChatGPT compared to the original human-written text. It provides a mechanism to measure the degree of ChatGPT influence in the resulting text. Our experimental results show our proposed model has better robustness on the HPPT dataset and two existing datasets (HC3 and CDB). Furthermore, the "Polish Ratio" we proposed offers a more comprehensive explanation by quantifying the degree of ChatGPT involvement.

📄 PDF Abstract BibTeX arXiv:2307.11380

Code (2)

clement1290/chatgpt-detection-pr-hppt 공식 구현 pytorch
freedomintelligence/chatgpt-detection-pr-hppt pytorch

Tasks

MisinformationText Generation

Similar Papers 제목 키워드 기반

Evaluation of vector embedding models in clustering of text documents

2019-09-01 · RANLP 2019 9 · Tomasz Walkowiak, Mateusz Gniewkowski

The paper presents an evaluation of word embedding models in clustering of texts in the Polish language. Authors verified six different embedding models, starting from widely used word2vec, across fastText with character…

Clustering

CHEAT: A Large-scale Dataset for Detecting ChatGPT-writtEn AbsTracts

2023-04-24 · Peipeng Yu, Jiahan Chen, Xuan Feng, Zhihua Xia

The powerful ability of ChatGPT has caused widespread concern in the academic community. Malicious users could synthesize dummy academic content through ChatGPT, which is extremely harmful to academic rigor and originali…

Do LLMs produce texts with "human-like" lexical diversity?

2025-07-31 · Kelly Kendro, Jeffrey Maloney, Scott Jarvis arxiv

The degree to which large language models (LLMs) produce writing that is truly human-like remains unclear despite the extensive empirical attention that this question has received. The present study addresses this questi…

Detection of Criminal Texts for the Polish State Border Guard

2021-08-24 · Artur Nowakowski, Krzysztof Jassem

This paper describes research on the detection of Polish criminal texts appearing on the Internet. We carried out experiments to find the best available setup for the efficient classification of unbalanced and noisy data…

Language ModelingLanguage Modelling

DiaBiz – an Annotated Corpus of Polish Call Center Dialogs

2022-06-01 · LREC 2022 6 · Piotr Pęzik, Gosia Krawentek, Sylwia Karasińska, Paweł Wilk 외

This paper introduces DiaBiz, a large, annotated, multimodal corpus of Polish telephone conversations conducted in varied business settings, comprising 4036 call centre interactions from nine different domains, i.e. bank…

speech-recognitionSpeech Recognition