paper-with-me

홈 › Papers

M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection

2024-02-17 · Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Osama Mohanned Afzal, Tarek Mahmoud, Giovanni Puccetti, Thomas Arnold, Alham Fikri Aji, Nizar Habash, Iryna Gurevych, Preslav Nakov

The advent of Large Language Models (LLMs) has brought an unprecedented surge in machine-generated text (MGT) across diverse channels. This raises legitimate concerns about its potential misuse and societal implications. The need to identify and differentiate such content from genuine human-generated text is critical in combating disinformation, preserving the integrity of education and scientific fields, and maintaining trust in communication. In this work, we address this problem by introducing a new benchmark based on a multilingual, multi-domain, and multi-generator corpus of MGTs -- M4GT-Bench. The benchmark is compiled of three tasks: (1) mono-lingual and multi-lingual binary MGT detection; (2) multi-way detection where one need to identify, which particular model generated the text; and (3) mixed human-machine text detection, where a word boundary delimiting MGT from human-written content should be determined. On the developed benchmark, we have tested several MGT detection baselines and also conducted an evaluation of human performance. We see that obtaining good performance in MGT detection usually requires an access to the training data from the same domain and generators. The benchmark is available at https://github.com/mbzuai-nlp/M4GT-Bench.

📄 PDF Abstract BibTeX arXiv:2402.11175

Code (1)

mbzuai-nlp/m4gt-bench 공식 구현

Tasks

Task 2Text Detection

Similar Papers 제목 키워드 기반

MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark

2023-10-20 · Dominik Macko, Robert Moro, Adaku Uchendu, Jason Samuel Lucas 외

There is a lack of research into capabilities of recent LLMs to generate convincing text in languages other than English and into performance of detectors of machine-generated text in multilingual settings. This is also …

Benchmarkingde-enText Detection

An Empirical Study of Explainable AI Techniques on Deep Learning Models For Time Series Tasks

2020-12-08 · Udo Schlegel, Daniela Oelke, Daniel A. Keim, Mennatallah El-Assady

Decision explanations of machine learning black-box models are often generated by applying Explainable AI (XAI) techniques. However, many proposed XAI methods produce unverified outputs. Evaluation and verification are u…

BIG-bench Machine LearningExplainable Artificial Intelligence (XAI)Time SeriesTime Series Analysis

DynamoRep: Trajectory-Based Population Dynamics for Classification of Black-box Optimization Problems

2023-06-08 · Gjorgjina Cenikj, Gašper Petelin, Carola Doerr, Peter Korošec 외

The application of machine learning (ML) models to the analysis of optimization algorithms requires the representation of optimization problems using numerical features. These features can be used as input for ML models …

BenchmarkingDescriptive

Landscape Analysis for Surrogate Models in the Evolutionary Black-Box Context

2022-02-11 · Zbyněk Pitra, Jan Koza, Jiří Tumpach, Martin Holeňa

Surrogate modeling has become a valuable technique for black-box optimization tasks with expensive evaluation of the objective function. In this paper, we investigate the relationship between the predictive accuracy of s…

TFRBench: A Reasoning Benchmark for Evaluating Forecasting Systems

2026-04-07 · Md Atik Ahamed, Mihir Parmar, Palash Goyal, Yiwen Song 외 arxiv

We introduce TFRBench, the first benchmark designed to evaluate the reasoning capabilities of forecasting systems. Traditionally, time-series forecasting has been evaluated solely on numerical accuracy, treating foundati…