paper-with-me

홈 › Papers

The Impact of Input Order Bias on Large Language Models for Software Fault Localization

2024-12-25 · Md Nakhla Rafi, Dong Jae Kim, Tse-Hsun Chen, Shaowei Wang

Large Language Models (LLMs) have shown significant potential in software engineering tasks such as Fault Localization (FL) and Automatic Program Repair (APR). This study investigates how input order and context size influence LLM performance in FL, a crucial step for many downstream software engineering tasks. We evaluate different method orderings using Kendall Tau distances, including "perfect" (where ground truths appear first) and "worst" (where ground truths appear last), across two benchmarks containing Java and Python projects. Our results reveal a strong order bias: in Java projects, Top-1 FL accuracy drops from 57% to 20% when reversing the order, while in Python projects, it decreases from 38% to approximately 3%. However, segmenting inputs into smaller contexts mitigates this bias, reducing the performance gap in FL from 22% and 6% to just 1% across both benchmarks. We replaced method names with semantically meaningful alternatives to determine whether this bias is due to data leakage. The observed trends remained consistent, suggesting that the bias is not caused by memorization from training data but rather by the inherent effect of input order. Additionally, we explored ordering methods based on traditional FL techniques and metrics, finding that DepGraph's ranking achieves 48% Top-1 accuracy, outperforming simpler approaches such as CallGraph(DFS). These findings highlight the importance of structuring inputs, managing context effectively, and selecting appropriate ordering strategies to enhance LLM performance in FL and other software engineering applications.

📄 PDF Abstract BibTeX arXiv:2412.18750

Code (0)

등록된 구현이 없습니다.

Tasks

Fault localizationMemorizationProgram Repair

Similar Papers 제목 키워드 기반

Unveiling Selection Biases: Exploring Order and Token Sensitivity in Large Language Models

2024-06-05 · Sheng-Lun Wei, Cheng-Kuang Wu, Hen-Hsen Huang, Hsin-Hsi Chen

In this paper, we investigate the phenomena of "selection biases" in Large Language Models (LLMs), focusing on problems where models are tasked with choosing the optimal option from an ordered sequence. We delve into bia…

Decision MakingSensitivity

What Kind of Language is Easy to Language-Model Under Curriculum Learning?

2026-04-29 · Nadine El-Naggar, Tatsuki Kuribayashi, Ted Briscoe arxiv

Many of the thousands of attested languages share common configurations of features, creating a spectrum from typologically very rare (e.g., object-verb-subject word order) or impossible languages to very common combinat…

Unexplored flaws in multiple-choice VQA evaluations

2025-11-27 · Fabio Rosenthal, Sebastian Schmidt, Thorsten Graf, Thorsten Bagodonat 외 arxiv

Multimodal Large Language Models (MLLMs) demonstrate strong capabilities in handling image-text inputs. A common way to assess this ability is through multiple-choice Visual Question Answering (VQA). Earlier works have a…

Visual Question Answering

On Positional Bias of Faithfulness for Long-form Summarization

2024-10-31 · David Wan, Jesse Vig, Mohit Bansal, Shafiq Joty

Large Language Models (LLMs) often exhibit positional bias in long-context settings, under-attending to information in the middle of inputs. We investigate the presence of this bias in long-form summarization, its impact…

Form

Capturing Failures of Large Language Models via Human Cognitive Biases

2022-02-24 · Erik Jones, Jacob Steinhardt

Large language models generate complex, open-ended outputs: instead of outputting a class label they write summaries, generate dialogue, or produce working code. In order to asses the reliability of these open-ended gene…

Code Generation