paper-with-me

Papers

MultiSocial: Multilingual Benchmark of Machine-Generated Text Detection of Social-Media Texts

2024-06-18 · Dominik Macko, Jakub Kopal, Robert Moro, Ivan Srba

Recent LLMs are able to generate high-quality multilingual texts, indistinguishable for humans from authentic human-written ones. Research in machine-generated text detection is however mostly focused on the English language and longer texts, such as news articles, scientific papers or student essays. Social-media texts are usually much shorter and often feature informal language, grammatical errors, or distinct linguistic items (e.g., emoticons, hashtags). There is a gap in studying the ability of existing methods in detection of such texts, reflected also in the lack of existing multilingual benchmark datasets. To fill this gap we propose the first multilingual (22 languages) and multi-platform (5 social media platforms) dataset for benchmarking machine-generated text detection in the social-media domain, called MultiSocial. It contains 472,097 texts, of which about 58k are human-written and approximately the same amount is generated by each of 7 multilingual LLMs. We use this benchmark to compare existing detection methods in zero-shot as well as fine-tuned form. Our results indicate that the fine-tuned detectors have no problem to be trained on social-media texts and that the platform selection for training matters.

📄 PDF Abstract BibTeX arXiv:2406.12549

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesBenchmarkingText Detection

Similar Papers 제목 키워드 기반

MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark

2023-10-20 · Dominik Macko, Robert Moro, Adaku Uchendu, Jason Samuel Lucas 외

There is a lack of research into capabilities of recent LLMs to generate convincing text in languages other than English and into performance of detectors of machine-generated text in multilingual settings. This is also …

Benchmarkingde-enText Detection

SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection

2024-04-22 · Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su 외

We present the results and the main findings of SemEval-2024 Task 8: Multigenerator, Multidomain, and Multilingual Machine-Generated Text Detection. The task featured three subtasks. Subtask A is a binary classification …

Binary ClassificationText Detection

Fine-tuning Large Language Models for Multigenerator, Multidomain, and Multilingual Machine-Generated Text Detection

2024-01-22 · Feng Xiong, Thanet Markchom, Ziwei Zheng, Subin Jung 외

SemEval-2024 Task 8 introduces the challenge of identifying machine-generated texts from diverse Large Language Models (LLMs) in various languages and domains. The task comprises three subtasks: binary classification in …

Binary ClassificationClassificationMulti-class Classificationtext-classification+2

CEAID: Benchmark of Multilingual Machine-Generated Text Detection Methods for Central European Languages

2025-09-30 · Dominik Macko, Jakub Kopal arxiv

Machine-generated text detection, as an important task, is predominantly focused on English in research. This makes the existing detectors almost unusable for non-English languages, relying purely on cross-lingual transf…

Adversarial RobustnessText Detection

LuxVeri at GenAI Detection Task 1: Inverse Perplexity Weighted Ensemble for Robust Detection of AI-Generated Text across English and Multilingual Contexts

2025-01-21 · Md Kamrujjaman Mobin, Md Saiful Islam

This paper presents a system developed for Task 1 of the COLING 2025 Workshop on Detecting AI-Generated Content, focusing on the binary classification of machine-generated versus human-written text. Our approach utilizes…

Binary ClassificationText Detection