paper-with-me

홈 › Papers

Linguistic and Embedding-Based Profiling of Texts generated by Humans and Large Language Models

2025-07-18 · Sergio E. Zanotto, Segun Aroyehun arxiv

The rapid advancements in large language models (LLMs) have significantly improved their ability to generate natural language, making texts generated by LLMs increasingly indistinguishable from human-written texts. While recent research has primarily focused on using LLMs to classify text as either human-written or machine-generated texts, our study focuses on characterizing these texts using a set of linguistic features across different linguistic levels such as morphology, syntax, and semantics. We select a dataset of human-written and machine-generated texts spanning 8 domains and produced by 11 different LLMs. We calculate different linguistic features such as dependency length and emotionality, and we use them for characterizing human-written and machine-generated texts along with different sampling strategies, repetition controls, and model release dates. Our statistical analysis reveals that human-written texts tend to exhibit simpler syntactic structures and more diverse semantic content. Furthermore, we calculate the variability of our set of features across models and domains. Both human- and machine-generated texts show stylistic diversity across domains, with human-written texts displaying greater variation in our features. Finally, we apply style embeddings to further test variability among human-written and machine-generated texts. Notably, newer models output text that is similarly variable, pointing to a homogenization of machine-generated texts.

📄 PDF Abstract BibTeX arXiv:2507.13614

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Human Variability vs. Machine Consistency: A Linguistic Analysis of Texts Generated by Humans and Large Language Models

2024-12-04 · Sergio E. Zanotto, Segun Aroyehun

The rapid advancements in large language models (LLMs) have significantly improved their ability to generate natural language, making texts generated by LLMs increasingly indistinguishable from human-written texts. Recen…

Semantic SimilaritySemantic Textual Similarity

Profiling-UD: a Tool for Linguistic Profiling of Texts

2020-05-01 · LREC 2020 5 · Dominique Brunato, Andrea Cimino, Felice Dell{'}Orletta, Giulia Venturi 외

In this paper, we introduce Profiling{--}UD, a new text analysis tool inspired to the principles of linguistic profiling that can support language variation research from different perspectives. It allows the extraction …

Author Profiling

UMUTextStats: A linguistic feature extraction tool for Spanish

2022-06-01 · LREC 2022 6 · José Antonio García-Díaz, Pedro José Vivancos-Vicente, Ángela Almela, Rafael Valencia-García

Feature Engineering consists in the application of domain knowledge to select and transform relevant features to build efficient machine learning models. In the Natural Language Processing field, the state of the art con…

Author ProfilingAuthorship VerificationDocument ClassificationFeature Engineering+3

Linguistic Profiling of Texts Across Textual Genres and Readability Levels. An Exploratory Study on Italian Fictional Prose

2013-09-01 · RANLP 2013 9 · Felice Dell{'}Orletta, Simonetta Montemagni, Giulia Venturi
Language IdentificationText Classification

Gender Prediction in English-Hindi Code-Mixed Social Media Content : Corpus and Baseline System

2018-06-14 · Ankush Khandelwal, Sahil Swami, Syed Sarfaraz Akhtar, Manish Shrivastava

The rapid expansion in the usage of social media networking sites leads to a huge amount of unprocessed user generated data which can be used for text mining. Author profiling is the problem of automatically determining …

Author ProfilingGender PredictionGeneral ClassificationLanguage Identification+3