paper-with-me

홈 › Papers

Isolating LLM Lexical Bias: A Curation-Free Triangulated Metric for Preference-Stage Learning

2026-05-29 · Xiaoyang Ming, Jose Hernandez, Thomas Stephan Juzek arxiv

Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models and their misalignment with natural language usage. These misalignments are thought to partly originate in the preference-learning stage, e.g. Reinforcement Learning from Human Feedback, which generally makes models more useful but simultaneously may introduce systematic lexical bias. In terms of lexical behavior, this is visible in a model's preference for certain formats or the overuse of words (delve, furthermore), even when such patterns are not present in base model outputs. Research on lexical misalignment induced during preference training is constrained by reliance on manual curation. We address this, by introducing the Triangulated Preference Shift score, a metric that triangulates between human gold standards, base models, and instruct variants to isolate shifts induced specifically by preference learning, without manual curation. We provide data across six model families, anchor the results in the literature, and illustrate the general approach's utility by analyzing whether preference learning shifts models toward what could be interpreted as a "language of prestige". The metric provides an initial automated method to quantify behavioral shifts attributable to preference tuning, and thus, may help inform model alignment and development of trustworthy AI.

📄 PDF Abstract BibTeX arXiv:2606.00334

Code (0)

등록된 구현이 없습니다.

Tasks

Reinforcement Learning

Similar Papers 제목 키워드 기반

Fully Automated Identification of Lexical Alignment and Preference-Stage Shifts in Large Language Models

2026-06-02 · Thomas Stephan Juzek, Xiaoyang Ming, Jose A. Hernandez arxiv

The language used by digital chat assistants such as ChatGPT can diverge from human expectations (misalignment). Research, mostly on Scientific English, has described both what divergences occur and, to some extent, why,…

A Technique for Isolating Lexically-Independent Phonetic Dependencies in Generative CNNs

2025-06-10 · Bruno Ferenc Šegedin

The ability of deep neural networks (DNNs) to represent phonotactic generalizations derived from lexical learning remains an open question. This study (1) investigates the lexically-invariant generalization capacity of g…

Computational Modeling of Affixoid Behavior in Chinese Morphology

2020-12-01 · COLING 2020 8 · Yu-Hsiang Tseng, Shu-Kai Hsieh, Pei-Yi Chen, Sara Court

The morphological status of affixes in Chinese has long been a matter of debate. How one might apply the conventional criteria of free/bound and content/function features to distinguish word-forming affixes from bound ro…

Instruction Tuning Chronologically Consistent Language Models

2025-10-13 · Songrun He, Linying Lv, Asaf Manela, Jimmy Wu arxiv

We introduce a family of chronologically consistent, instruction-tuned large language models to eliminate lookahead bias. Each model is trained only on data available before a clearly defined knowledge-cutoff date, ensur…

NumtaDB - Assembled Bengali Handwritten Digits

2018-06-06 · Samiul Alam, Tahsin Reasat, Rashed Mohammad Doha, Ahmed Imtiaz Humayun

To benchmark Bengali digit recognition algorithms, a large publicly available dataset is required which is free from biases originating from geographical location, gender, and age. With this aim in mind, NumtaDB, a datas…