paper-with-me

홈 › Papers

Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs

2025-09-01 · Andong Hua, Kenan Tang, Chenhe Gu, Jindong Gu, Eric Wong, Yao Qin arxiv

Prompt sensitivity, referring to the phenomenon where paraphrasing (i.e., repeating something written or spoken using different words) leads to significant changes in large language model (LLM) performance, has been widely accepted as a core limitation of LLMs. In this work, we revisit this issue and ask: Is the widely reported high prompt sensitivity truly an inherent weakness of LLMs, or is it largely an artifact of evaluation processes? To answer this question, we systematically evaluate 7 LLMs (e.g., GPT and Gemini family) across 6 benchmarks, including both multiple-choice and open-ended tasks on 12 diverse prompt templates. We find that much of the prompt sensitivity stems from heuristic evaluation methods, including log-likelihood scoring and rigid answer matching, which often overlook semantically correct responses expressed through alternative phrasings, such as synonyms or paraphrases. When we adopt LLM-as-a-Judge evaluations, we observe a substantial reduction in performance variance and a consistently higher correlation in model rankings across prompts. Our findings suggest that modern LLMs are more robust to prompt templates than previously believed, and that prompt sensitivity may be more an artifact of evaluation than a flaw in the models.

📄 PDF Abstract BibTeX arXiv:2509.01790

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval

2026-06-09 · João Maria Janeiro, Mathurin Videau, Andrea Caciolai, Benjamin Piwowarski 외 arxiv

Multiple-choice (MCQA) benchmarks are the standard for evaluating pretrained large language models, but their reliance on log-likelihood scoring makes them unreliable. Specifically, standard scores are highly sensitive t…

Rethinking Prompt-based Debiasing in Large Language Models

2025-03-12 · Xinyi Yang, Runzhe Zhan, Derek F. Wong, Shu Yang 외

Investigating bias in large language models (LLMs) is crucial for developing trustworthy AI. While prompt-based through prompt engineering is common, its effectiveness relies on the assumption that models inherently unde…

Prompt Engineering

OmniAID: Decoupling Semantic and Artifacts for Universal AI-Generated Image Detection in the Wild

2025-11-11 · Yuncheng Guo, Junyan Ye, Chenjue Zhang, Hengrui Kang 외 arxiv

A truly universal AI-Generated Image (AIGI) detector must simultaneously generalize across diverse generative models and varied semantic content. Current methods learn a single, entangled forgery representation, conflati…

Refining Visual Artifacts in Diffusion Models via Explainable AI-based Flaw Activation Maps

2025-12-09 · Seoyeon Lee, Gwangyeol Yu, Chaewon Kim, Jonghyuk Park arxiv

Diffusion models have achieved remarkable success in image synthesis. However, addressing artifacts and unrealistic regions remains a critical challenge. We propose self-refining diffusion, a novel framework that enhance…

Text-to-Image Generation

Evaluating Remote Sensing Image Captions Beyond Metric Biases

2026-04-22 · Ziyun Chen, Fan Liu, Liang Yao, Chuanyi Zhang 외 arxiv

The core objective of image captioning is to achieve lossless semantic compression from visual signals into textual modalities. However, the reliance on manually curated reference texts for evaluation essentially forces …

Image Captioning