paper-with-me

홈 › Papers

Adaptive Rigor in AI System Evaluation using Temperature-Controlled Verdict Aggregation via Generalized Power Mean

2026-04-04 · Aleksandr Meshkov arxiv

Existing evaluation methods for LLM-based AI systems, such as LLM-as-a-Judge, verdict systems, and NLI, do not always align well with human assessment because they cannot adapt their strictness to the application domain. This paper presents Temperature-Controlled Verdict Aggregation (TCVA), a method that combines a five-level verdict-scoring system with generalized power-mean aggregation and an intuitive temperature parameter T [0.1, 1.0] to control evaluation rigor. Low temperatures yield pessimistic scores suited for safety-critical domains; high temperatures produce lenient scores appropriate for conversational AI. Experimental evaluation on three benchmark datasets with human Likert-scale annotations (SummEval and USR) shows that TCVA achieves correlation with human judgments comparable to RAGAS on faithfulness (Spearman = 0.667 vs. 0.676) while consistently outperforming DeepEval. The method requires no additional LLM calls when adjusting the temperature parameter.

📄 PDF Abstract BibTeX arXiv:2604.08595

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Infinite-Dimensional Adaptive Boundary Observer for Inner-Domain Temperature Estimation of 3D Electrosurgical Processes using Surface Thermography Sensing

2022-11-01 · Hamza El-Kebir, Junren Ran, Martin Ostoja-Starzewski, Richard Berlin 외

We present a novel 3D adaptive observer framework for use in the determination of subsurface organic tissue temperatures in electrosurgery. The observer structure leverages pointwise 2D surface temperature readings obtai…

parameter estimationTime SeriesTime Series Analysis

Residual-Conservative Model Predictive Path Integral Control

2026-07-08 · Hyung-Jin Yoon, Hunmin Kim arxiv

Sampling-based model predictive control methods handle nonlinear dynamics and complex cost landscapes through Monte Carlo rollouts, yet typically employ fixed constraint penalties that do not adapt to model-plant mismatc…

CIVET: Systematic Evaluation of Understanding in VLMs

2025-06-05 · Massimo Rizzoli, Simone Alghisi, Olha Khomyn, Gabriel Roccabruna 외

While Vision-Language Models (VLMs) have achieved competitive performance in various tasks, their comprehension of the underlying structure and semantics of a scene remains understudied. To investigate the understanding …

Object

Soft Adaptive Policy Optimization

2025-11-25 · Chang Gao, Chujie Zheng, Xiong-Hui Chen, Kai Dang 외 arxiv

Reinforcement learning (RL) plays an increasingly important role in enhancing the reasoning capabilities of large language models (LLMs), yet stable and performant policy optimization remains challenging. Token-level imp…

Reinforcement LearningMathematical Reasoning

Optimization of Temperature and Relative Humidity in an Automatic Egg Incubator Using Mamdani Interference System

2022-06-17 · Pramit Dutta, Nafisa Anjum

Temperature and humidity are two of the rudimentary factors that must be controlled during egg incubation. Improper temperature and humidity levels during the incubation period often result in unwanted conditions. This p…