paper-with-me

홈 › Papers

LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs

2024-08-16 · Do Xuan Long, Hai Nguyen Ngoc, Tiviatis Sim, Hieu Dao, Shafiq Joty, Kenji Kawaguchi, Nancy F. Chen, Min-Yen Kan

We present the first systematic evaluation examining format bias in performance of large language models (LLMs). Our approach distinguishes between two categories of an evaluation metric under format constraints to reliably and accurately assess performance: one measures performance when format constraints are adhered to, while the other evaluates performance regardless of constraint adherence. We then define a metric for measuring the format bias of LLMs and establish effective strategies to reduce it. Subsequently, we present our empirical format bias evaluation spanning four commonly used categories -- multiple-choice question-answer, wrapping, list, and mapping -- covering 15 widely-used formats. Our evaluation on eight generation tasks uncovers significant format bias across state-of-the-art LLMs. We further discover that improving the format-instruction following capabilities of LLMs across formats potentially reduces format bias. Based on our evaluation findings, we study prompting and fine-tuning with synthesized format data techniques to mitigate format bias. Our methods successfully reduce the variance in ChatGPT's performance among wrapping formats from 235.33 to 0.71 (%$^2$).

📄 PDF Abstract BibTeX arXiv:2408.08656

Code (1)

dxlong2000/FormatEval 공식 구현

Tasks

Instruction FollowingMultiple-choice

Similar Papers 제목 키워드 기반

StructEval: Benchmarking LLMs' Capabilities to Generate Structural Outputs

2025-05-26 · Jialin Yang, Dongfu Jiang, Lipeng He, Sherman Siu 외

As Large Language Models (LLMs) become integral to software development workflows, their ability to generate structured outputs has become critically important. We introduce StructEval, a comprehensive benchmark for eval…

Benchmarking

Evaluating and Mitigating Social Bias for Large Language Models in Open-ended Settings

2024-12-09 · Zhao Liu, Tian Xie, Xueru Zhang

Current social bias benchmarks for Large Language Models (LLMs) primarily rely on pre-defined question formats like multiple-choice, limiting their ability to reflect the complexity and open-ended nature of real-world in…

Multiple-choice

Are LLMs Ready for TOON? Benchmarking Structural Correctness-Sustainability Trade-offs in Novel Structured Output Formats

2026-01-17 · Elio Masciari, Vincenzo Moscato, Enea Vincenzo Napolitano, Gian Marco Orlando 외 arxiv

Large Language Models (LLMs) are increasingly required to generate structured, machine-readable outputs for downstream systems. While recent benchmarks have focused on evaluating the structural correctness of such output…

Benchmarking Bias in Large Language Models during Role-Playing

2024-11-01 · Xinyue Li, Zhenpeng Chen, Jie M. Zhang, Yiling Lou 외

Large Language Models (LLMs) have become foundational in modern language-driven applications, profoundly influencing daily life. A critical technique in leveraging their potential is role-playing, where LLMs simulate div…

BenchmarkingFairnessMultiple-choice

Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement

2026-06-16 · Ramaravind Kommiya Mothilal, Terry Jingchen Zhang, Raiyan Ahmed, Zhijing Jin 외 arxiv

Evaluations of social bias in LLMs largely focus on whether models generate or imply biased content. However, as LLMs are increasingly used as judges of bias, they may exhibit social biases in subtler ways in how they ev…

Logical Reasoning