paper-with-me

홈 › Papers

Do Large Language Models Rank Fairly? An Empirical Study on the Fairness of LLMs as Rankers

2024-04-04 · YuAn Wang, Xuyang Wu, Hsin-Tai Wu, Zhiqiang Tao, Yi Fang

The integration of Large Language Models (LLMs) in information retrieval has raised a critical reevaluation of fairness in the text-ranking models. LLMs, such as GPT models and Llama2, have shown effectiveness in natural language understanding tasks, and prior works (e.g., RankGPT) have also demonstrated that the LLMs exhibit better performance than the traditional ranking models in the ranking task. However, their fairness remains largely unexplored. This paper presents an empirical study evaluating these LLMs using the TREC Fair Ranking dataset, focusing on the representation of binary protected attributes such as gender and geographic location, which are historically underrepresented in search outcomes. Our analysis delves into how these LLMs handle queries and documents related to these attributes, aiming to uncover biases in their ranking algorithms. We assess fairness from both user and content perspectives, contributing an empirical benchmark for evaluating LLMs as the fair ranker.

📄 PDF Abstract BibTeX arXiv:2404.03192

Code (0)

등록된 구현이 없습니다.

Tasks

FairnessInformation RetrievalNatural Language UnderstandingRetrieval

Methods 이 논문이 사용한 방법론

Refunds@Expedia|||How do I get a full refund from Expedia? “How do I get a full refund from Expedia? How do I get a full refund from Expedia? – Call ☎️ +1-(888) 829 (0881) or +1-805-330-4056 or +1-805-330-4056 for Quick Help &…
Attention 설명 없음
Linear Layer A Linear Layer is a projection $\mathbf{XW + b}$.
Residual Connection 설명 없음
Dropout Dropout is a regularization technique for neural networks that drops a unit (along with connections) at training time with a specified probability $p$ (a common value is…
Layer Normalization Unlike batch normalization, Layer Normalization directly estimates the normalization statistics from the summed inputs…
Multi-Head Attention 설명 없음
Weight Decay 설명 없음

Similar Papers 제목 키워드 기반

Subverting Fair Image Search with Generative Adversarial Perturbations

2022-05-05 · Avijit Ghosh, Matthew Jagielski, Christo Wilson

In this work we explore the intersection fairness and robustness in the context of ranking: when a ranking model has been calibrated to achieve some definition of fairness, is it possible for an external adversary to mak…

FairnessImage RetrievalRe-Ranking

Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks

2024-06-12 · Justin Zhao, Flor Miriam Plaza-del-Arco, Benjie Genchel, Amanda Cercas Curry

As Large Language Models (LLMs) continue to evolve, the search for efficient and meaningful evaluation methods is ongoing. Many recent evaluations use LLMs as judges to score outputs from other LLMs, often relying on a s…

BenchmarkingChatbotEmotional IntelligenceLanguage Modeling+2

A Distributed Frank-Wolfe Algorithm for Communication-Efficient Sparse Learning

2014-04-09 · Aurélien Bellet, YIngyu Liang, Alireza Bagheri Garakani, Maria-Florina Balcan 외

Learning sparse combinations is a frequent theme in machine learning. In this paper, we study its associated optimization problem in the distributed setting where the elements to be combined are not centrally located but…

Sparse Learning

An Empirical Analysis of GPT-4V's Performance on Fashion Aesthetic Evaluation

2024-10-31 · Yuki Hirakawa, Takashi Wada, Kazuya Morishita, Ryotaro Shimizu 외

Fashion aesthetic evaluation is the task of estimating how well the outfits worn by individuals in images suit them. In this work, we examine the zero-shot performance of GPT-4V on this task for the first time. We show t…

Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation

2025-05-22 · Jiwon Moon, Yerin Hwang, Dongryeol Lee, Taegwan Kang 외

With the growing use of large language models(LLMs) as evaluators, their application has expanded to code evaluation tasks, where they assess the correctness of generated code without relying on reference implementations…