paper-with-me

홈 › Papers

Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judge

2025-05-25 · Zhuo Liu, Moxin Li, Xun Deng, Qifan Wang, Fuli Feng

LLM-as-a-Judge employs large language models (LLMs), such as GPT-4, to evaluate the quality of LLM-generated responses, gaining popularity for its cost-effectiveness and strong alignment with human evaluations. However, training proxy judge models using evaluation data generated by powerful teacher models introduces a critical yet previously overlooked issue: teacher preference bias, where the proxy judge model learns a biased preference for responses from the teacher model. To tackle this problem, we propose a novel setting that incorporates an additional assistant model, which is not biased toward the teacher model's responses, to complement the training data. Building on this setup, we introduce AGDe-Judge, a three-stage framework designed to debias from both the labels and feedbacks in the training data. Extensive experiments demonstrate that AGDe-Judge effectively reduces teacher preference bias while maintaining strong performance across six evaluation benchmarks. Code is available at https://github.com/Liuz233/AGDe-Judge.

📄 PDF Abstract BibTeX arXiv:2505.19176

Code (1)

liuz233/agde-judge 공식 구현

Methods 이 논문이 사용한 방법론

Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Position-Wise Feed-Forward Layer 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Label Smoothing Label Smoothing is a regularization technique that introduces noise for the labels. This accounts for the fact that datasets may have mistakes in them, so maximizing the…
Multi-Head Attention 설명 없음

Similar Papers 제목 키워드 기반

Pareto-Guided Teacher Alignment for Fair Personalized Text Generation

2026-06-08 · Tunazzina Islam arxiv

Personalized persuasive text generation can improve relevance and engagement, but demographic conditioning may also introduce unequal framing across groups. We study fairness mitigation in personalized generation as a co…

Text Generation

Densely Guided Knowledge Distillation using Multiple Teacher Assistants

2020-09-18 · ICCV 2021 10 · Wonchul Son, Jaemin Na, Junyong Choi, Wonjun Hwang

With the success of deep neural networks, knowledge distillation which guides the learning of a small student network from a large teacher network is being actively studied for model compression and transfer learning. Ho…

Knowledge DistillationModel CompressionTransfer Learning

Who Gets Cited? Gender- and Majority-Bias in LLM-Driven Reference Selection

2025-08-02 · Jiangen He arxiv

Large language models (LLMs) are rapidly being adopted as research assistants, particularly for literature review and reference recommendation, yet little is known about whether they introduce demographic bias into citat…

Mitigating LLM biases toward spurious social contexts using direct preference optimization

2026-04-02 · Hyunji Nam, Dorottya Demszky arxiv

LLMs are increasingly used for high-stakes decision-making, yet their sensitivity to spurious contextual information can introduce harmful biases. This is a critical concern when models are deployed for tasks like evalua…

User-Assistant Bias in LLMs

2025-08-16 · Xu Pan, Jingxuan Fan, Zidi Xiong, Ely Hahami 외 arxiv

Modern large language models (LLMs) are typically trained and deployed using structured role tags (e.g. system, user, assistant, tool) that explicitly mark the source of each piece of context. While these tags are essent…

Instruction Following