paper-with-me

홈 › Papers

Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback

2025-08-03 · Tom S. Juzek, Zina B. Ward arxiv

Large Language Models (LLMs) are known to overuse certain terms like "delve" and "intricate." The exact reasons for these lexical choices, however, have been unclear. Using Meta's Llama model, this study investigates the contribution of Learning from Human Feedback (LHF), under which we subsume Reinforcement Learning from Human Feedback and Direct Preference Optimization. We present a straightforward procedure for detecting the lexical preferences of LLMs that are potentially LHF-induced. Next, we more conclusively link LHF to lexical overuse by experimentally emulating the LHF procedure and demonstrating that participants systematically prefer text variants that include certain words. This lexical overuse can be seen as a sort of misalignment, though our study highlights the potential divergence between the lexical expectations of different populations -- namely LHF workers versus LLM users. Our work contributes to the growing body of research on explainable artificial intelligence and emphasizes the importance of both data and procedural transparency in alignment research.

📄 PDF Abstract BibTeX arXiv:2508.01930

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Position: The Term "Machine Unlearning" Is Overused in LLMs

2026-05-08 · Sangyeon Yoon, Yeachan Jun, Albert No arxiv

Large language models increasingly face demands to "forget" training data, knowledge, or behaviors due to regulatory deletion obligations, copyright/licensing disputes, and safety or product-policy requirements. This pos…

Why Does ChatGPT "Delve" So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models

2024-12-16 · Tom S. Juzek, Zina B. Ward

Scientific English is currently undergoing rapid change, with words like "delve," "intricate," and "underscore" appearing far more frequently than just a few years ago. It is widely assumed that scientists' use of large …

Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models

2026-06-02 · Thomas Stephan Juzek, Xiaoyang Ming, Jose A. Hernandez arxiv

The language used by digital chat assistants such as ChatGPT can diverge from human expectations (misalignment). Research, mostly on Scientific English, has described both what divergences occur and, to some extent, why,…

Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning

2026-05-29 · Xiaoyang Ming, Jose Hernandez, Thomas Stephan Juzek arxiv

Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models and their misalignment with natural language usage. These misalignmen…

Reinforcement Learning

Unified Lexical Representation for Interpretable Visual-Language Alignment

2024-07-25 · YiFan Li, Yikai Wang, Yanwei Fu, Dongyu Ru 외

Visual-Language Alignment (VLA) has gained a lot of attention since CLIP's groundbreaking work. Although CLIP performs well, the typical direct latent feature alignment lacks clarity in its representation and similarity …

Cross-Modal RetrievalLanguage Modelling