paper-with-me

Papers

Benchmarking Large Language Models via Random Variables

2025-01-20 · Zijin Hong, Hao Wu, Su Dong, Junnan Dong, Yilin Xiao, Yujing Zhang, Zhu Wang, Feiran Huang, Linyi Li, Hongxia Yang, Xiao Huang

With the continuous advancement of large language models (LLMs) in mathematical reasoning, evaluating their performance in this domain has become a prominent research focus. Recent studies have raised concerns about the reliability of current mathematical benchmarks, highlighting issues such as simplistic design and potential data leakage. Therefore, creating a reliable benchmark that effectively evaluates the genuine capabilities of LLMs in mathematical reasoning remains a significant challenge. To address this, we propose RV-Bench, a framework for Benchmarking LLMs via Random Variables in mathematical reasoning. Specifically, the background content of a random variable question (RV question) mirrors the original problem in existing standard benchmarks, but the variable combinations are randomized into different values. LLMs must fully understand the problem-solving process for the original problem to correctly answer RV questions with various combinations of variable values. As a result, the LLM's genuine capability in mathematical reasoning is reflected by its accuracy on RV-Bench. Extensive experiments are conducted with 29 representative LLMs across 900+ RV questions. A leaderboard for RV-Bench ranks the genuine capability of these LLMs. Further analysis of accuracy dropping indicates that current LLMs still struggle with complex mathematical reasoning problems.

📄 PDF Abstract BibTeX arXiv:2501.11790

Code (0)

등록된 구현이 없습니다.

Tasks

BenchmarkingMathematical Reasoning

Similar Papers 제목 키워드 기반

Risk Aware Benchmarking of Large Language Models

2023-10-11 · Apoorva Nitsure, Youssef Mroueh, Mattia Rigotti, Kristjan Greenewald 외

We propose a distributional framework for benchmarking socio-technical risks of foundation models with quantified statistical significance. Our approach hinges on a new statistical relative testing based on first and sec…

BenchmarkingEconometricsModel SelectionPortfolio Optimization

Intrinsic uncertainties and where to find them

2021-07-06 · Francesco Farina, Lawrence Phillips, Nicola J Richmond

We introduce a framework for uncertainty estimation that both describes and extends many existing methods. We consider typical hyperparameters involved in classical training as random variables and marginalise them out t…

Benchmarking

An Interpretable Measure for Quantifying Predictive Dependence between Continuous Random Variables -- Extended Version

2025-01-18 · Renato Assunção, Flávio Figueiredo, Francisco N. Tinoco Júnior, Léo M. de Sá-Freire 외

A fundamental task in statistical learning is quantifying the joint dependence or association between two continuous random variables. We introduce a novel, fully non-parametric measure that assesses the degree of associ…

Benchmarking

Synthetic location trajectory generation using categorical diffusion models

2024-02-19 · Simon Dirmeier, Ye Hong, Fernando Perez-Cruz

Diffusion probabilistic models (DPMs) have rapidly evolved to be one of the predominant generative models for the simulation of synthetic data, for instance, for computer vision, audio, natural language processing, or bi…

BenchmarkingDecision MakingSynthetic Data Generation

Benchmarking Machine Learning Models to Predict Corporate Bankruptcy

2022-12-22 · Emmanuel Alanis, Sudheer Chava, Agam Shah

Using a comprehensive sample of 2,585 bankruptcies from 1990 to 2019, we benchmark the performance of various machine learning models in predicting financial distress of publicly traded U.S. firms. We find that gradient …

Benchmarking