paper-with-me

Papers

MULTITuDE: Large-Scale Multilingual Machine-Generated Text Detection Benchmark

2023-10-20 · Dominik Macko, Robert Moro, Adaku Uchendu, Jason Samuel Lucas, Michiharu Yamashita, Matúš Pikuliak, Ivan Srba, Thai Le, Dongwon Lee, Jakub Simko, Maria Bielikova

There is a lack of research into capabilities of recent LLMs to generate convincing text in languages other than English and into performance of detectors of machine-generated text in multilingual settings. This is also reflected in the available benchmarks which lack authentic texts in languages other than English and predominantly cover older generators. To fill this gap, we introduce MULTITuDE, a novel benchmarking dataset for multilingual machine-generated text detection comprising of 74,081 authentic and machine-generated texts in 11 languages (ar, ca, cs, de, en, es, nl, pt, ru, uk, and zh) generated by 8 multilingual LLMs. Using this benchmark, we compare the performance of zero-shot (statistical and black-box) and fine-tuned detectors. Considering the multilinguality, we evaluate 1) how these detectors generalize to unseen languages (linguistically similar as well as dissimilar) and unseen LLMs and 2) whether the detectors improve their performance when trained on multiple languages.

📄 PDF Abstract BibTeX arXiv:2310.13606

Code (1)

kinit-sk/mgt-detection-benchmark 공식 구현 pytorch

Tasks

Benchmarkingde-enText Detection

Similar Papers 제목 키워드 기반

Polyglots or Multitudes? Multilingual LLM Answers to Value-laden Multiple-Choice Questions

2026-02-05 · Léo Labat, Etienne Ollion, François Yvon arxiv

Multiple-Choice Questions (MCQs) are often used to assess knowledge, reasoning abilities, and even values encoded in large language models (LLMs). While the effect of multilingualism has been studied on LLM factual recal…

Machine Translation

Fineweb-Edu-Ar: Machine-translated Corpus to Support Arabic Small Language Models

2024-11-10 · Sultan Alrashed, Dmitrii Khizbullin, David R. Pugh

As large language models (LLMs) grow and develop, so do their data demands. This is especially true for multilingual LLMs, where the scarcity of high-quality and readily available data online has led to a multitude of sy…

Dataset GenerationMachine Translation

Template-based multilingual football reports generation using Wikidata as a knowledge base

2018-11-01 · WS 2018 11 · Lorenzo Gatti, Chris van der Lee, Mari{\"e}t Theune

This paper presents a new version of a football reports generation system called PASS. The original version generated Dutch text and relied on a limited hand-crafted knowledge base. We describe how, in a short amount of …

Machine TranslationText GenerationTranslation

SemEval-2024 Task 8: Multidomain, Multimodel and Multilingual Machine-Generated Text Detection

2024-04-22 · Yuxia Wang, Jonibek Mansurov, Petar Ivanov, Jinyan Su 외

We present the results and the main findings of SemEval-2024 Task 8: Multigenerator, Multidomain, and Multilingual Machine-Generated Text Detection. The task featured three subtasks. Subtask A is a binary classification …

Binary ClassificationText Detection

Fine-tuning Large Language Models for Multigenerator, Multidomain, and Multilingual Machine-Generated Text Detection

2024-01-22 · Feng Xiong, Thanet Markchom, Ziwei Zheng, Subin Jung 외

SemEval-2024 Task 8 introduces the challenge of identifying machine-generated texts from diverse Large Language Models (LLMs) in various languages and domains. The task comprises three subtasks: binary classification in …

Binary ClassificationClassificationMulti-class Classificationtext-classification+2