Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering
Evaluating answers from state-of-the-art large language models (LLMs) is challenging: lexical metrics miss semantic nuances, whereas "LLM-as-Judge" scoring is computationally expensive. We re-evaluate a lightweight alternative -- off-the-shelf Natural Language Inference (NLI) scoring augmented by a simple lexical-match flag and find that this decades-old technique matches GPT-4o's accuracy (89.9%) on long-form QA, while requiring orders-of-magnitude fewer parameters. To test human alignment of these metrics rigorously, we introduce DIVER-QA, a new 3000-sample human-annotated benchmark spanning five QA datasets and five candidate LLMs. Our results highlight that inexpensive NLI-based evaluation remains competitive and offer DIVER-QA as an open resource for future metric research.
Code (0)
등록된 구현이 없습니다.
Tasks
Natural Language InferenceQuestion AnsweringSimilar Papers 제목 키워드 기반
NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference Checklist
In this study, we analyze automatic evaluation metrics for Natural Language Generation (NLG), specifically task-agnostic metrics and human-aligned metrics. Task-agnostic metrics, such as Perplexity, BLEU, BERTScore, are …
Controllable Language ModellingDialogue GenerationLanguage Modellingnlg evaluation+3Metrics vs Surveys: An Analysis for Human-Aligned Benchmarking in Social Robot Navigation
Social, also called human-aware, navigation is a key challenge for integrating mobile robots into human environments. The evaluation of such systems is complex, as factors such as comfort, safety, and legibility must be …
Robot NavigationRevisiting Training-free NAS Metrics: An Efficient Training-based Method
Recent neural architecture search (NAS) works proposed training-free metrics to rank networks which largely reduced the search cost in NAS. In this paper, we revisit these training-free metrics and find that: (1) the num…
GPUNeural Architecture SearchRevisiting Misalignment in Multispectral Pedestrian Detection: A Language-Driven Approach for Cross-modal Alignment Fusion
Multispectral pedestrian detection is a crucial component in various critical applications. However, a significant challenge arises due to the misalignment between these modalities, particularly under real-world conditio…
cross-modal alignmentPedestrian DetectionRevisiting Automatic Question Summarization Evaluation in the Biomedical Domain
Automatic evaluation metrics have been facilitating the rapid development of automatic summarization methods by providing instant and fair assessments of the quality of summaries. Most metrics have been developed for the…
Text Generation