paper-with-me

홈 › Papers

OpenTuringBench: An Open-Model-based Benchmark and Framework for Machine-Generated Text Detection and Attribution

2025-04-15 · Lucio La Cava, Andrea Tagarelli

Open Large Language Models (OLLMs) are increasingly leveraged in generative AI applications, posing new challenges for detecting their outputs. We propose OpenTuringBench, a new benchmark based on OLLMs, designed to train and evaluate machine-generated text detectors on the Turing Test and Authorship Attribution problems. OpenTuringBench focuses on a representative set of OLLMs, and features a number of challenging evaluation tasks, including human/machine-manipulated texts, out-of-domain texts, and texts from previously unseen models. We also provide OTBDetector, a contrastive learning framework to detect and attribute OLLM-based machine-generated texts. Results highlight the relevance and varying degrees of difficulty of the OpenTuringBench tasks, with our detector achieving remarkable capabilities across the various tasks and outperforming most existing detectors. Resources are available on the OpenTuringBench Hugging Face repository at https://huggingface.co/datasets/MLNTeam-Unical/OpenTuringBench

📄 PDF Abstract BibTeX arXiv:2504.11369

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeAuthorship AttributionContrastive LearningText Detection

Methods 이 논문이 사용한 방법론

Contrastive Learning 설명 없음
SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

OpenLS-DGF: An Adaptive Open-Source Dataset Generation Framework for Machine Learning Tasks in Logic Synthesis

2024-11-14 · Liwei Ni, Rui Wang, Miao Liu, Xingyu Meng 외

This paper introduces OpenLS-DGF, an adaptive logic synthesis dataset generation framework, to enhance machine learning~(ML) applications within the logic synthesis process. Previous dataset generation flows were tailore…

Dataset Generation

RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors

2024-05-13 · Liam Dugan, Alyssa Hwang, Filip Trhlik, Josh Magnus Ludan 외

Many commercial and open-source models claim to detect machine-generated text with extremely high accuracy (99% or more). However, very few of these detectors are evaluated on shared benchmark datasets and even when they…

Adversarial RobustnessText Detection

Attribution Quality in AI-Generated Content:Benchmarking Style Embeddings and LLM Judges

2025-10-14 · Misam Abbas arxiv

Attributing authorship in the era of large language models (LLMs) is increasingly challenging as machine-generated prose rivals human writing. We benchmark two complementary attribution mechanisms , fixed Style Embedding…

FinML-Chain: A Blockchain-Integrated Dataset for Enhanced Financial Machine Learning

2024-11-25 · Jingfeng Chen, Wanlin Deng, Dangxing Chen, Luyao Zhang

Machine learning is critical for innovation and efficiency in financial markets, offering predictive models and data-driven decision-making. However, challenges such as missing data, lack of transparency, untimely update…

Fairness

The Limitations of Stylometry for Detecting Machine-Generated Fake News

2019-08-26 · CL 2020 6 · Tal Schuster, Roei Schuster, Darsh J Shah, Regina Barzilay

Recent developments in neural language models (LMs) have raised concerns about their potential misuse for automatically spreading misinformation. In light of these concerns, several studies have proposed to detect machin…

Fake News DetectionLanguage ModellingMisinformation