paper-with-me

홈 › Papers

AI-Associated Lexical Shifts Across 34 Languages: Cross-Lingual Convergence and Diachronic Uptake in News Writing

2026-05-25 · Thomas Stephan Juzek arxiv

AI-associated lexical shifts have been documented mainly in Scientific English. We extend this work to 34 languages in the WMT News Crawl corpus, refining a split-halves continuation diagnostic that compares GPT-4.1 continuations with matched human gold-standard text. For each language, we derive ranked AI-overused lemmas using log prevalence ratios. We find substantial cross-lingual semantic convergence: semantically related concepts recur across typologically diverse languages, with 'emphasize'-type verbs appearing in 24 of 34 languages. Embedding-based and manual analyses support this pattern. We also examine diachronic uptake in news writing before and after ChatGPT's release. Tracking each language's top 20 AI-overused items, we find prevalence increases in 26 of 34 languages from 2020-2021 to 2023-2024, with a mean change of +15.1%, whilst matched baseline words show no comparable increase (-4.5%). In 10 languages with longer historical coverage, longitudinal analyses show post-2022 increases that exceed the modest shifts observed in earlier periods, though with smaller effect sizes than in Scientific English. We validate our approach extensively, including across seeds, model variants, data sizes, model families, and more. Our findings are consistent with the view that AI-associated lexical preferences extend beyond English and may exert cross-lingual homogenising pressure on global language use.

📄 PDF Abstract BibTeX arXiv:2605.25358

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Delexicalized Word Embeddings for Cross-lingual Dependency Parsing

2017-04-01 · EACL 2017 4 · Mathieu Dehouck, Pascal Denis

This paper presents a new approach to the problem of cross-lingual dependency parsing, aiming at leveraging training data from different source languages to learn a parser in a target language. Specifically, this approac…

Cross-Lingual TransferDependency ParsingWord Embeddings

Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent

2025-05-12 · Ethan Gotlieb Wilcox, Cui Ding, Giovanni Acampa, Tiago Pimentel 외

This paper argues that the relationship between lexical identity and prosody -- one well-studied parameter of linguistic variation -- can be characterized using information theory. We predict that languages that use pros…

Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models

2026-06-02 · Thomas Stephan Juzek, Xiaoyang Ming, Jose A. Hernandez arxiv

The language used by digital chat assistants such as ChatGPT can diverge from human expectations (misalignment). Research, mostly on Scientific English, has described both what divergences occur and, to some extent, why,…

Survey of Computational Approaches to Lexical Semantic Change

2018-11-15 · Nina Tahmasebi, Lars Borin, Adam Jatowt

Our languages are in constant flux driven by external factors such as cultural, societal and technological changes, as well as by only partially understood internal motivations. Words acquire new meanings and lose old se…

Change DetectionInformation RetrievalOptical Character Recognition (OCR)Retrieval+1

Measuring Lexical Similarity across Sign Languages in Global Signbank

2020-05-01 · LREC 2020 5 · Carl B{\"o}rstell, Onno Crasborn, Lori Whynot

Lexicostatistics is the main method used in previous work measuring linguistic distances between sign languages. As a method, it disregards any possible structural/grammatical similarity, instead focusing exclusively on …