paper-with-me

홈 › Papers

Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models

2026-06-02 · Thomas Stephan Juzek, Xiaoyang Ming, Jose A. Hernandez arxiv

The language used by digital chat assistants such as ChatGPT can diverge from human expectations (misalignment). Research, mostly on Scientific English, has described both what divergences occur and, to some extent, why, linking them to the training stage of human preference learning. Yet, existing approaches rely on manual curation. This paper introduces two curation-free, assumption-light evaluation metrics: the Lexical Alignment Score, which identifies lexical overuse, and the Triangulated Preference Shift, which quantifies how much of such shifts can be attributed to human preference learning. Using PubMed abstracts, continuations were generated and measured using windowed document prevalence across six model families (Falcon, Gemma, Llama, Mistral, OLMo, Yi). The procedure identifies, without manual intervention, overused items such as 'suggest', 'additionally', and 'strategy', and estimates their link to preference learning. Our findings replicate prior work and remain stable across parameter settings, random seeds, and evaluation on further data. The approach scales readily and enables systematic study of lexical (mis)alignment beyond Scientific English and across languages, and as such, the metrics have the potential to contribute to improved alignment for future models and understanding of its origins.

📄 PDF Abstract BibTeX arXiv:2606.03165

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning

2026-05-29 · Xiaoyang Ming, Jose Hernandez, Thomas Stephan Juzek arxiv

Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models and their misalignment with natural language usage. These misalignmen…

Reinforcement Learning

Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback

2025-08-03 · Tom S. Juzek, Zina B. Ward arxiv

Large Language Models (LLMs) are known to overuse certain terms like "delve" and "intricate." The exact reasons for these lexical choices, however, have been unclear. Using Meta's Llama model, this study investigates the…

Reinforcement Learning

Is Preference Alignment Always the Best Option to Enhance LLM-Based Translation? An Empirical Analysis

2024-09-30 · Hippolyte Gisserot-Boukhlef, Ricardo Rei, Emmanuel Malherbe, Céline Hudelot 외

Neural metrics for machine translation (MT) evaluation have become increasingly prominent due to their superior correlation with human judgments compared to traditional lexical metrics. Researchers have therefore utilize…

Machine TranslationTranslation

The CQC Algorithm: Cycling in Graphs to Semantically Enrich and Enhance a Bilingual Dictionary

2014-01-18 · Tiziano Flati, Roberto Navigli

Bilingual machine-readable dictionaries are knowledge resources useful in many automatic tasks. However, compared to monolingual computational lexicons like WordNet, bilingual dictionaries typically provide a lower amoun…

TAGTranslation

Can Safety Emerge from Weak Supervision? A Systematic Analysis of Small Language Models

2026-03-07 · Punyajoy Saha, Sudipta Halder, Debjyoti Mondal, Subhadarshi Panda arxiv

Safety alignment is critical for deploying large language models (LLMs) in real-world applications, yet most existing approaches rely on large human-annotated datasets and static red-teaming benchmarks that are costly, d…