paper-with-me

홈 › Papers

Classification of Human- and AI-Generated Texts for English, French, German, and Spanish

2023-12-08 · Kristina Schaaff, Tim Schlippe, Lorenz Mindner

In this paper we analyze features to classify human- and AI-generated text for English, French, German and Spanish and compare them across languages. We investigate two scenarios: (1) The detection of text generated by AI from scratch, and (2) the detection of text rephrased by AI. For training and testing the classifiers in this multilingual setting, we created a new text corpus covering 10 topics for each language. For the detection of AI-generated text, the combination of all proposed features performs best, indicating that our features are portable to other related languages: The F1-scores are close with 99% for Spanish, 98% for English, 97% for German and 95% for French. For the detection of AI-rephrased text, the systems with all features outperform systems with other features in many cases, but using only document features performs best for German (72%) and Spanish (86%) and only text vector features leads to best results for English (78%).

📄 PDF Abstract BibTeX arXiv:2312.04882

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Visual Cues and Error Correction for Translation Robustness

2021-03-12 · Findings (EMNLP) 2021 11 · Zhenhao Li, Marek Rei, Lucia Specia

Neural Machine Translation models are sensitive to noise in the input texts, such as misspelled words and ungrammatical constructions. Existing robustness techniques generally fail when faced with unseen types of noise a…

Machine TranslationTranslation

Can AI-Generated Persuasion Be Detected? Persuaficial Benchmark and AI vs. Human Linguistic Differences

2026-01-08 · Arkadiusz Modzelewski, Paweł Golik, Anna Kołos, Giovanni Da San Martino arxiv

Large Language Models (LLMs) can generate highly persuasive text, raising concerns about their misuse for propaganda, manipulation, and other harmful purposes. This leads us to our central question: Is LLM-generated pers…

Deriving Disinformation Insights from Geolocalized Twitter Callouts

2021-08-06 · David Tuxworth, Dimosthenis Antypas, Luis Espinosa-Anke, Jose Camacho-Collados 외

This paper demonstrates a two-stage method for deriving insights from social media data relating to disinformation by applying a combination of geospatial classification and embedding-based language modelling across mult…

Language ModellingSpecificityWord Embeddings

A Data-driven Approach to Named Entity Recognition for Early Modern French

2022-10-01 · COLING 2022 10 · Pedro Ortiz Suarez, Simon Gabay

Named entity recognition has become an increasingly useful tool for digital humanities research, specially when it comes to historical texts. However, historical texts pose a wide range of challenges to both named entity…

named-entity-recognitionNamed Entity RecognitionNamed Entity Recognition (NER)NER

On the differences between human translations

2020-11-01 · EAMT 2020 11 · Maja Popovic

Many studies have confirmed that translated texts exhibit different features than texts originally written in the given language. This work explores texts translated by different translators taking into account expertise…

Machine TranslationSentenceTranslation