paper-with-me

Papers

How NOT to benchmark your SITE metric: Beyond Static Leaderboards and Towards Realistic Evaluation

2025-10-07 · Prabhant Singh, Sibylle Hess, Joaquin Vanschoren arxiv

Transferability estimation metrics are used to find a high-performing pre-trained model for a given target task without fine-tuning models and without access to the source dataset. Despite the growing interest in developing such metrics, the benchmarks used to measure their progress have gone largely unexamined. In this work, we empirically show the shortcomings of widely used benchmark setups to evaluate transferability estimation metrics. We argue that the benchmarks on which these metrics are evaluated are fundamentally flawed. We empirically demonstrate that their unrealistic model spaces and static performance hierarchies artificially inflate the perceived performance of existing metrics, to the point where simple, dataset-agnostic heuristics can outperform sophisticated methods. Our analysis reveals a critical disconnect between current evaluation protocols and the complexities of real-world model selection. To address this, we provide concrete recommendations for constructing more robust and realistic benchmarks to guide future research in a more meaningful direction.

📄 PDF Abstract BibTeX arXiv:2510.06448

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

ClawBench: Can AI Agents Complete Everyday Online Tasks?

2026-04-09 · Yuxuan Zhang, Yubo Wang, Yipeng Zhu, Penghui Du 외 arxiv

AI agents may be able to automate your inbox, but can they automate other routine aspects of your life? Everyday online tasks offer a realistic yet unsolved testbed for evaluating the next generation of AI agents. To thi…

Hire Your Anthropologist! Rethinking Culture Benchmarks Through an Anthropological Lens

2025-10-07 · Mai AlKhamissi, Yunze Xiao, Badr AlKhamissi, Mona Diab arxiv

Cultural evaluation of large language models has become increasingly important, yet current benchmarks often reduce culture to static facts or homogeneous values. This view conflicts with anthropological accounts that em…

YourBench: Easy Custom Evaluation Sets for Everyone

2025-04-02 · Sumuk Shashidhar, Clémentine Fourrier, Alina Lozovskia, Thomas Wolf 외

Evaluating large language models (LLMs) effectively remains a critical bottleneck, as traditional static benchmarks suffer from saturation and contamination, while human evaluations are costly and slow. This hinders time…

MMLU

We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

2024-07-01 · Runqi Qiao, Qiuna Tan, Guanting Dong, Minhui Wu 외

Visual mathematical reasoning, as a fundamental visual reasoning ability, has received widespread attention from the Large Multimodal Models (LMMs) community. Existing benchmarks, such as MathVista and MathVerse, focus m…

MathMathematical ReasoningMemorizationVisual Reasoning

Muizalix Pro

2025-02-07 · 02/07 2025 2 · Muizalix Pro

At Muizalix, we’re a dedicated team of creative professionals and digital experts passionate about helping businesses grow. Specializing in social media marketing, SEO, content creation, and website design, we provide co…

Marketing