paper-with-me

홈 › Papers

A Systematic Analysis of Linguistic Features in AI-Generated Text Detection Across Domains and Models

2026-06-02 · Yassir El Attar, Esra Dönmez, Maximilian Maurer, Agnieszka Falenska arxiv

Interpretable linguistic features offer a promising approach for explaining why a given text appears machine-generated, particularly for non-expert users. However, existing findings on which features reliably indicate LLM-generated text remain fragmented across feature sets, models, and text domains. To address this gap, we conduct a large-scale empirical study assessing the robustness of linguistic signals for characterizing AI-generated text. Our analysis covers 284 interpretable linguistic features across outputs from 27 LLMs and ten text domains under cross-model and cross-domain generalization settings. We show that classifiers based solely on linguistic features can reliably distinguish AI-generated from human-written text. However, many previously proposed indicators prove strongly context-dependent, with the exception of measures of lexical richness, which remain robust signals across model families and text domains. These results demonstrate which linguistic signals generalize across contexts and provide a foundation for more reliable, interpretable analyses of AI-generated language.

📄 PDF Abstract BibTeX arXiv:2606.04177

Code (0)

등록된 구현이 없습니다.

Tasks

Domain GeneralizationText Detection

Similar Papers 제목 키워드 기반

Long-form analogies generated by chatGPT lack human-like psycholinguistic properties

2023-06-07 · S. M. Seals, Valerie L. Shalin

Psycholinguistic analyses provide a means of evaluating large language model (LLM) output and making systematic comparisons to human-generated text. These methods can be used to characterize the psycholinguistic properti…

FormLanguage ModelingLanguage ModellingLarge Language Model

Explaining Generalization of AI-Generated Text Detectors Through Linguistic Analysis

2026-01-12 · Yuxi Xia, Kinga Stańczak, Benjamin Roth arxiv

AI-text detectors achieve high accuracy on in-domain benchmarks, but often struggle to generalize across different generation conditions such as unseen prompts, model families, or domains. While prior work has reported t…

Differentiating between human-written and AI-generated texts using linguistic features automatically extracted from an online computational tool

2024-07-04 · Georgios P. Georgiou

While extensive research has focused on ChatGPT in recent years, very few studies have systematically quantified and compared linguistic features between human-written and Artificial Intelligence (AI)-generated language.…

Human Variability vs. Machine Consistency: A Linguistic Analysis of Texts Generated by Humans and Large Language Models

2024-12-04 · Sergio E. Zanotto, Segun Aroyehun

The rapid advancements in large language models (LLMs) have significantly improved their ability to generate natural language, making texts generated by LLMs increasingly indistinguishable from human-written texts. Recen…

Semantic SimilaritySemantic Textual Similarity

Linguistic and Embedding-Based Profiling of Texts generated by Humans and Large Language Models

2025-07-18 · Sergio E. Zanotto, Segun Aroyehun arxiv

The rapid advancements in large language models (LLMs) have significantly improved their ability to generate natural language, making texts generated by LLMs increasingly indistinguishable from human-written texts. While…