paper-with-me

홈 › Papers

Which LLMs are Difficult to Detect? A Detailed Analysis of Potential Factors Contributing to Difficulties in LLM Text Detection

2024-10-18 · Shantanu Thorat, Tianbao Yang

As LLMs increase in accessibility, LLM-generated texts have proliferated across several fields, such as scientific, academic, and creative writing. However, LLMs are not created equally; they may have different architectures and training datasets. Thus, some LLMs may be more challenging to detect than others. Using two datasets spanning four total writing domains, we train AI-generated (AIG) text classifiers using the LibAUC library - a deep learning library for training classifiers with imbalanced datasets. Our results in the Deepfake Text dataset show that AIG-text detection varies across domains, with scientific writing being relatively challenging. In the Rewritten Ivy Panda (RIP) dataset focusing on student essays, we find that the OpenAI family of LLMs was substantially difficult for our classifiers to distinguish from human texts. Additionally, we explore possible factors that could explain the difficulties in detecting OpenAI-generated texts.

📄 PDF Abstract BibTeX arXiv:2410.14875

Code (1)

ShantanuT01/which-llms-are-difficult-to-detect 공식 구현 pytorch

Tasks

Text Detection

Methods 이 논문이 사용한 방법론

Library 설명 없음

Similar Papers 제목 키워드 기반

An Empirical Analysis of Factual Errors in Human-Written Text and its Application

2026-06-26 · Kazuma Iwamoto, Kazumasa Omura, Shotaro Ishihara arxiv

Factual Error Detection (FED), which is the task of identifying factually incorrect spans in a given text, has long been recognized as an important research problem. However, with the rapid rise of large language models …

VulDetectBench: Evaluating the Deep Capability of Vulnerability Detection with Large Language Models

2024-06-11 · Yu Liu, Lang Gao, Mingxin Yang, Yu Xie 외

Large Language Models (LLMs) have training corpora containing large amounts of program code, greatly improving the model's code comprehension and generation capabilities. However, sound comprehensive research on detectin…

Vulnerability Detection

DEEM: Dynamic Experienced Expert Modeling for Stance Detection

2024-02-23 · Xiaolong Wang, Yile Wang, Sijie Cheng, Peng Li 외

Recent work has made a preliminary attempt to use large language models (LLMs) to solve the stance detection task, showing promising results. However, considering that stance detection usually requires detailed backgroun…

Stance Detection

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage

2024-12-20 · Saehyung Lee, Seunghyun Yoon, Trung Bui, Jing Shi 외

Multimodal large language models (MLLMs) excel at generating highly detailed captions but often produce hallucinations. Our analysis reveals that existing hallucination detection methods struggle with detailed captions. …

AttributeBenchmarkingHallucinationImage Captioning+1

From Calculation to Adjudication: Examining LLM judges on Mathematical Reasoning Tasks

2024-09-06 · Andreas Stephan, Dawei Zhu, Matthias Aßenmacher, Xiaoyu Shen 외

To reduce the need for human annotations, large language models (LLMs) have been proposed as judges of the quality of other candidate models. The performance of LLM judges is typically evaluated by measuring the correlat…

Machine TranslationMathematical Reasoning