paper-with-me

홈 › Papers

Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering

2025-11-10 · Sai Shridhar Balamurali, Lu Cheng arxiv

Evaluating answers from state-of-the-art large language models (LLMs) is challenging: lexical metrics miss semantic nuances, whereas "LLM-as-Judge" scoring is computationally expensive. We re-evaluate a lightweight alternative -- off-the-shelf Natural Language Inference (NLI) scoring augmented by a simple lexical-match flag and find that this decades-old technique matches GPT-4o's accuracy (89.9%) on long-form QA, while requiring orders-of-magnitude fewer parameters. To test human alignment of these metrics rigorously, we introduce DIVER-QA, a new 3000-sample human-annotated benchmark spanning five QA datasets and five candidate LLMs. Our results highlight that inexpensive NLI-based evaluation remains competitive and offer DIVER-QA as an open resource for future metric research.

📄 PDF Abstract BibTeX arXiv:2511.07659

Code (0)

등록된 구현이 없습니다.

Tasks

Natural Language InferenceQuestion Answering

Similar Papers 제목 키워드 기반

NLG Evaluation Metrics Beyond Correlation Analysis: An Empirical Metric Preference Checklist

2023-05-15 · Iftitahu Ni'mah, Meng Fang, Vlado Menkovski, Mykola Pechenizkiy

In this study, we analyze automatic evaluation metrics for Natural Language Generation (NLG), specifically task-agnostic metrics and human-aligned metrics. Task-agnostic metrics, such as Perplexity, BLEU, BERTScore, are …

Controllable Language ModellingDialogue GenerationLanguage Modellingnlg evaluation+3

Metrics vs Surveys: An Analysis for Human-Aligned Benchmarking in Social Robot Navigation

2025-10-03 · Stefano Trepella, Mauro Martini, Noé Pérez-Higueras, Andrea Ostuni 외 arxiv

Social, also called human-aware, navigation is a key challenge for integrating mobile robots into human environments. The evaluation of such systems is complex, as factors such as comfort, safety, and legibility must be …

Robot Navigation

Revisiting Training-free NAS Metrics: An Efficient Training-based Method

2022-11-16 · Taojiannan Yang, Linjie Yang, Xiaojie Jin, Chen Chen

Recent neural architecture search (NAS) works proposed training-free metrics to rank networks which largely reduced the search cost in NAS. In this paper, we revisit these training-free metrics and find that: (1) the num…

GPUNeural Architecture Search

Revisiting Misalignment in Multispectral Pedestrian Detection: A Language-Driven Approach for Cross-modal Alignment Fusion

2024-11-27 · Taeheon Kim, Sangyun Chung, Youngjoon Yu, Yong Man Ro

Multispectral pedestrian detection is a crucial component in various critical applications. However, a significant challenge arises due to the misalignment between these modalities, particularly under real-world conditio…

cross-modal alignmentPedestrian Detection

Revisiting Automatic Question Summarization Evaluation in the Biomedical Domain

2023-03-18 · Hongyi Yuan, Yaoyun Zhang, Fei Huang, Songfang Huang

Automatic evaluation metrics have been facilitating the rapid development of automatic summarization methods by providing instant and fair assessments of the quality of summaries. Most metrics have been developed for the…

Text Generation