paper-with-me

Papers

Fine-Grained Benchmark Generation for Comprehensive Evaluation of Foundation Models

2026-05-12 · Mohammed Saidul Islam, Negin Baghbanzadeh, Farnaz Kohankhaki, Afshin Cheraghi, Ali Kore, Shayaan Mehdi, Elham Dolatabadi, Arash Afkanpour arxiv

Evaluation of foundation models often rely on aggregate scores from benchmarks that lack comprehensive coverage and metadata for a fine-grained evaluation. We introduce a framework for automated benchmark generation. Our framework generates evaluation problems grounded in reference material, such as textbooks, producing benchmarks with broad coverage, rich metadata, and robustness to contamination. The pipeline employs a multi-agent architecture for problem generation and a solution-graph-driven strategy that significantly improves the reliability of ground truth solutions. Using the framework, we generate three benchmarks in Machine Learning, Corporate Finance, and Personal Finance. Expert review finds a significantly lower ground-truth error rate than previous benchmarks such as MMLU and GSM8K. Evaluation of 12 commercial and open-source models shows that our benchmarks achieve near-uniform competency coverage and surface performance differences across models that existing benchmarks fail to capture. We will open-source the framework and our curated benchmarks soon.

📄 PDF Abstract BibTeX arXiv:2605.18824

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EvalMuse-40K: A Reliable and Fine-Grained Benchmark with Comprehensive Human Annotations for Text-to-Image Generation Model Evaluation

2024-12-24 · Shuhao Han, Haotian Fan, Jiachen Fu, Liang Li 외

Recently, Text-to-Image (T2I) generation models have achieved significant advancements. Correspondingly, many automated metrics have emerged to evaluate the image-text alignment capabilities of generative models. However…

Image CaptioningImage GenerationText to Image GenerationText-to-Image Generation+1

FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video Generation

2023-11-03 · NeurIPS 2023 11 · Yuanxin Liu, Lei LI, Shuhuai Ren, Rundong Gao 외

Recently, open-domain text-to-video (T2V) generation models have made remarkable progress. However, the promising results are mainly shown by the qualitative cases of generated videos, while the quantitative evaluation o…

Text-to-Video GenerationVideo Generation

AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation

2026-04-09 · Ziwei Zhou, Zeyuan Lai, Rui Wang, Yifan Yang 외 arxiv

Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and video in isolation or rely on coarse embe…

Video Generation

Fine-grained Claim-level RAG Benchmark for Law

2026-05-20 · Souvick Das, Sallam Abualhaija, Domenico Bianculli arxiv

The rapid progress of large language models (LLMs) is shifting semantic search toward a question-answering paradigm, where users ask questions and LLMs generate responses. In high-stake domains such as law, retrieval-aug…

CT-FineBench: A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report Generation

2026-04-27 · Ruifeng Yuan, Wanxing Chang, Weiwei Cao, Bowen Shi 외 arxiv

The evaluation of generated reports remains a critical challenge in Computed Tomography (CT) report generation, due to the large volume of text, the diversity and complexity of findings, and the presence of fine-grained,…