paper-with-me

Papers

DAFE: LLM-Based Evaluation Through Dynamic Arbitration for Free-Form Question-Answering

2025-03-11 · Sher Badshah, Hassan Sajjad

Evaluating Large Language Models (LLMs) free-form generated responses remains a challenge due to their diverse and open-ended nature. Traditional supervised signal-based automatic metrics fail to capture semantic equivalence or handle the variability of open-ended responses, while human evaluation, though reliable, is resource-intensive. Leveraging LLMs as evaluators offers a promising alternative due to their strong language understanding and instruction-following capabilities. Taking advantage of these capabilities, we propose the Dynamic Arbitration Framework for Evaluation (DAFE), which employs two primary LLM-as-judges and engages a third arbitrator only in cases of disagreements. This selective arbitration prioritizes evaluation reliability while reducing unnecessary computational demands compared to conventional majority voting. DAFE utilizes task-specific reference answers with dynamic arbitration to enhance judgment accuracy, resulting in significant improvements in evaluation metrics such as Macro F1 and Cohen's Kappa. Through experiments, including a comprehensive human evaluation, we demonstrate DAFE's ability to provide consistent, scalable, and resource-efficient assessments, establishing it as a robust framework for evaluating free-form model outputs.

📄 PDF Abstract BibTeX arXiv:2503.08542

Code (0)

등록된 구현이 없습니다.

Tasks

FormInstruction FollowingQuestion Answering

Similar Papers 제목 키워드 기반

Value-of-Information based Arbitration between Model-based and Model-free Control

2019-12-08 · Krishn Bera, Yash Mandilwar, Bapi Raju

There have been numerous attempts in explaining the general learning behaviours using model-based and model-free methods. While the model-based control is flexible yet computationally expensive in planning, the model-fre…

Computational EfficiencymodelQ-LearningReinforcement Learning

AdaFed: Fair Federated Learning via Adaptive Common Descent Direction

2024-01-10 · Shayan Mohajer Hamidi, En-hui Yang

Federated learning (FL) is a promising technology via which some edge devices/clients collaboratively train a machine learning model orchestrated by a server. Learning an unfair model is known as a critical problem in fe…

Federated Learning

Accelerating Fair Federated Learning: Adaptive Federated Adam

2023-01-23 · Li Ju, Tianru Zhang, Salman Toor, Andreas Hellander

Federated learning is a distributed and privacy-preserving approach to train a statistical model collaboratively from decentralized data of different parties. However, when datasets of participants are not independent an…

FairnessFederated LearningPrivacy Preserving

Training-Free Generalized Few-Shot Segmentation through Open-Vocabulary Semantic Arbitration

2026-06-08 · Silas Kwabla Gah, Ebenezer Owusu arxiv

Generalized Few-Shot Semantic Segmentation (GFSS) has traditionally been approached as a representation-learning problem, requiring task-specific adaptation to incorporate novel classes from limited support examples. Rec…

Generalized Few-Shot Semantic Segmentation

SDAFE: A Dual-filter Stable Diffusion Data Augmentation Method for Facial Expression Recognition

2025-04-06 · ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2025 4 · Minghao Zhao, Yifei Chen, Jiahao Lyu, Shuangli Du 외

Facial expressions are a powerful medium for conveying emotions. In facial expression recognition (FER) field, the difficulty of collecting specific expressions often leads to class imbalance in mainstream datasets, sign…

Data AugmentationFacial Expression RecognitionFacial Expression Recognition (FER)