paper-with-me

Papers

MILE-RefHumEval: A Reference-Free, Multi-Independent LLM Framework for Human-Aligned Evaluation

2026-02-10 · Nalin Srun, Parisa Rastin, Guénaël Cabanes, Lydia Boudjeloud Assala arxiv

We introduce MILE-RefHumEval, a reference-free framework for evaluating Large Language Models (LLMs) without ground-truth annotations or evaluator coordination. It leverages an ensemble of independently prompted evaluators guided by a human-aligned schema, supporting both discrete and continuous scoring judgement. With task-specific prompts from best candidate selection, summarization and image captioning to dialogue, MILE-RefHumEval provides flexible, interpretable, and scalable assessments. Experiments show it aligns closely with human judgments, outperforms prior methods, and reduces computational overhead, offering an efficient, robust, and human-aligned solution for real-world LLM evaluation.

📄 PDF Abstract BibTeX arXiv:2602.09624

Code (0)

등록된 구현이 없습니다.

Tasks

Image Captioning

Similar Papers 제목 키워드 기반

GEN: Highly Efficient SMILES Explorer Using Autodidactic Generative Examination Networks

2019-09-10 · Ruud van Deursen, Peter Ertl, Igor V. Tetko, Guillaume Godin

Recurrent neural networks have been widely used to generate millions of de novo molecules in a known chemical space. These deep generative models are typically setup with LSTM or GRU units and trained with canonical SMIL…

valid

"Pale as death" or "pâle comme la mort" : Frozen similes used as literary clichés

2015-11-05 · Suzanne Mpouli, Jean-Gabriel Ganascia

The present study is focused on the automatic identification and description of frozen similes in British and French novels written between the 19 th century and the beginning of the 20 th century. Two main patterns of f…

Smiles in delta

2022-09-01 · Arianna Mingone

Fukasawa introduced in [Fukasawa, Math Financ, 2012] two necessary conditions for no butterfly arbitrage which require that the $d_1$ and $d_2$ functions of the Black-Scholes formula have to be decreasing. In this articl…

Math

Middle-mile logistics through the lens of goal-conditioned reinforcement learning

2026-05-04 · Onno Eberhard, Thibaut Cuvelier, Michal Valko, Bruno De Backer arxiv

Middle-mile logistics describes the problem of routing parcels through a network of hubs linked by trucks with finite capacity. We rephrase this as a multi-object goal-conditioned MDP. Our method combines graph neural ne…

Reinforcement Learning

Open Problem: Properly learning decision trees in polynomial time?

2022-06-29 · Guy Blanc, Jane Lange, Mingda Qiao, Li-Yang Tan

The authors recently gave an $n^{O(\log\log n)}$ time membership query algorithm for properly learning decision trees under the uniform distribution (Blanc et al., 2021). The previous fastest algorithm for this problem r…