paper-with-me

Papers

Fluid Language Model Benchmarking

2025-09-14 · Valentin Hofmann, David Heineman, Ian Magnusson, Kyle Lo, Jesse Dodge, Maarten Sap, Pang Wei Koh, Chun Wang, Hannaneh Hajishirzi, Noah A. Smith arxiv

Language model (LM) benchmarking faces several challenges: comprehensive evaluations are costly, benchmarks often fail to measure the intended capabilities, and evaluation quality can degrade due to labeling errors and benchmark saturation. Although various strategies have been proposed to mitigate these issues, they tend to address individual aspects in isolation, neglecting broader questions about overall evaluation quality. Here, we introduce Fluid Benchmarking, a new evaluation approach that advances LM benchmarking across multiple dimensions. Inspired by psychometrics, Fluid Benchmarking is based on the insight that the relative value of benchmark items depends on an LM's capability level, suggesting that evaluation should adapt to each LM. Methodologically, Fluid Benchmarking estimates an item response model based on existing LM evaluation results and uses the inferred quantities to select evaluation items dynamically, similar to computerized adaptive testing in education. In our experiments, we compare Fluid Benchmarking against the common practice of random item sampling as well as more sophisticated baselines, including alternative methods grounded in item response theory. We examine four dimensions -- efficiency, validity, variance, and saturation -- and find that Fluid Benchmarking achieves superior performance in all of them (e.g., higher validity and less variance on MMLU with fifty times fewer items). Our analysis shows that the two components of Fluid Benchmarking have distinct effects: item response theory, used to map performance into a latent ability space, increases validity, while dynamic item selection reduces variance. Overall, our results suggest that LM benchmarking can be substantially improved by moving beyond static evaluation.

📄 PDF Abstract BibTeX arXiv:2509.11106

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FluidLab: A Differentiable Environment for Benchmarking Complex Fluid Manipulation

2023-03-04 · Zhou Xian, Bo Zhu, Zhenjia Xu, Hsiao-Yu Tung 외

Humans manipulate various kinds of fluids in their everyday life: creating latte art, scooping floating objects from water, rolling an ice cream cone, etc. Using robots to augment or replace human labors in these daily s…

BenchmarkingGPU

Plug-and-Play Benchmarking of Reinforcement Learning Algorithms for Large-Scale Flow Control

2026-01-21 · Jannis Becktepe, Aleksandra Franz, Nils Thuerey, Sebastian Peitz arxiv

Reinforcement learning (RL) has shown promising results in active flow control (AFC), yet progress in the field remains difficult to assess as existing studies rely on heterogeneous observation and actuation schemes, num…

Reinforcement Learning

Benchmarking machine learning models for predicting aerofoil performance

2025-04-22 · Oliver Summerell, Gerardo Aragon-Camarasa, Stephanie Ordonez Sanchez

This paper investigates the capability of Neural Networks (NNs) as alternatives to the traditional methods to analyse the performance of aerofoils used in the wind and tidal energy industry. The current methods used to a…

Benchmarking

Geometry Matters: Benchmarking Scientific ML Approaches for Flow Prediction around Complex Geometries

2024-12-31 · Ali Rabeh, Ethan Herron, Aditya Balu, Soumik Sarkar 외

Rapid and accurate simulations of fluid dynamics around complicated geometric bodies are critical in a variety of engineering and scientific applications, including aerodynamics and biomedical flows. However, while scien…

BenchmarkingOut-of-Distribution Generalization

Benchmarking YOLOv5 and YOLOv7 models with DeepSORT for droplet tracking applications

2023-01-19 · Mihir Durve, Sibilla Orsini, Adriano Tiribocchi, Andrea Montessori 외

Tracking droplets in microfluidics is a challenging task. The difficulty arises in choosing a tool to analyze general microfluidic videos to infer physical quantities. The state-of-the-art object detector algorithm You O…

BenchmarkingGPUObject Tracking