paper-with-me

홈 › Papers

M4: Multi-generator, Multi-domain, and Multi-lingual Black-Box Machine-Generated Text Detection

2023-05-24 · Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su, Artem Shelmanov, Akim Tsvigun, Chenxi Whitehouse, Osama Mohammed Afzal, Tarek Mahmoud, Toru Sasaki, Thomas Arnold, Alham Fikri Aji, Nizar Habash, Iryna Gurevych, Preslav Nakov

Large language models (LLMs) have demonstrated remarkable capability to generate fluent responses to a wide variety of user queries. However, this has also raised concerns about the potential misuse of such texts in journalism, education, and academia. In this study, we strive to create automated systems that can detect machine-generated texts and pinpoint potential misuse. We first introduce a large-scale benchmark \textbf{M4}, which is a multi-generator, multi-domain, and multi-lingual corpus for machine-generated text detection. Through an extensive empirical study of this dataset, we show that it is challenging for detectors to generalize well on instances from unseen domains or LLMs. In such cases, detectors tend to misclassify machine-generated text as human-written. These results show that the problem is far from solved and that there is a lot of room for improvement. We believe that our dataset will enable future research towards more robust approaches to this pressing societal problem. The dataset is available at https://github.com/mbzuai-nlp/M4.

📄 PDF Abstract BibTeX arXiv:2305.14902

Code (2)

mbzuai-nlp/m4 공식 구현
mbzuai-nlp/semeval2024-task8 공식 구현

Tasks

Text Detection

Similar Papers 제목 키워드 기반

SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection

2024-04-22 · Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su 외

We present the results and the main findings of SemEval-2024 Task 8: Multigenerator, Multidomain, and Multilingual Machine-Generated Text Detection. The task featured three subtasks. Subtask A is a binary classification …

Binary ClassificationText Detection

Authorship Attribution in Multilingual Machine-Generated Texts

2025-08-03 · Lucio La Cava, Dominik Macko, Róbert Móro, Ivan Srba 외 arxiv

As Large Language Models (LLMs) have reached human-like fluency and coherence, distinguishing machine-generated text (MGT) from human-written content becomes increasingly difficult. While early efforts in MGT detection h…

Binary Classification

Fine-tuning Large Language Models for Multigenerator, Multidomain, and Multilingual Machine-Generated Text Detection

2024-01-22 · Feng Xiong, Thanet Markchom, Ziwei Zheng, Subin Jung 외

SemEval-2024 Task 8 introduces the challenge of identifying machine-generated texts from diverse Large Language Models (LLMs) in various languages and domains. The task comprises three subtasks: binary classification in …

Binary ClassificationClassificationMulti-class Classificationtext-classification+2

PetKaz at SemEval-2024 Task 8: Can Linguistics Capture the Specifics of LLM-generated Text?

2024-04-08 · Kseniia Petukhova, Roman Kazakov, Ekaterina Kochmar

In this paper, we present our submission to the SemEval-2024 Task 8 "Multigenerator, Multidomain, and Multilingual Black-Box Machine-Generated Text Detection", focusing on the detection of machine-generated texts (MGTs) …

DiversityText Detection

M4GT-Bench: Evaluation Benchmark for Black-Box Machine-Generated Text Detection

2024-02-17 · Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su 외

The advent of Large Language Models (LLMs) has brought an unprecedented surge in machine-generated text (MGT) across diverse channels. This raises legitimate concerns about its potential misuse and societal implications.…

Task 2Text Detection