paper-with-me

홈 › Papers

ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge

2025-10-21 · Zhilin Wang, Jaehun Jung, Ximing Lu, Shizhe Diao, Ellie Evans, Jiaqi Zeng, Pavlo Molchanov, Yejin Choi, Jan Kautz, Yi Dong arxiv

Evaluating progress in large language models (LLMs) is often constrained by the challenge of verifying responses, limiting assessments to tasks like mathematics, programming, and short-form question-answering. However, many real-world applications require evaluating LLMs in processing professional documents, synthesizing information, and generating comprehensive reports in response to user queries. We introduce ProfBench: a set of over 7000 response-criterion pairs as evaluated by human-experts with professional knowledge across Physics PhD, Chemistry PhD, Finance MBA and Consulting MBA. We build robust and affordable LLM-Judges to evaluate ProfBench rubrics, by mitigating self-enhancement bias and reducing the cost of evaluation by 2-3 orders of magnitude, to make it fair and accessible to the broader community. Our findings reveal that ProfBench poses significant challenges even for state-of-the-art LLMs, with top-performing models like GPT-5-high achieving only 65.9% overall performance. Furthermore, we identify notable performance disparities between proprietary and open-weight models and provide insights into the role that extended thinking plays in addressing complex, professional-domain tasks. Data: https://huggingface.co/datasets/nvidia/ProfBench and Code: https://github.com/NVlabs/ProfBench and Leaderboard: https://huggingface.co/spaces/nvidia/ProfBench

📄 PDF Abstract BibTeX arXiv:2510.18941

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks

2026-04-07 · Michael Krumdick, Varshini Reddy, Shivani Chaudhary, William Day 외 arxiv

As concerns surrounding AI-driven labor displacement intensify in knowledge-intensive sectors, existing benchmarks fail to measure performance on tasks that define practical professional expertise. Finance, in particular…

APTER: Adaptive Post-Training with Expert-Grounded Rubrics

2026-08-14 · Xukai Wang, Liangqi Li, Zhiyue Xu, Jingang Zhou 외 arxiv

As large language models enter professional domains, they must satisfy domain constraints, include critical evidence, and provide complete reasoning rather than merely produce fluent responses. Existing post-training met…

Mathematical ReasoningReinforcement LearningQuestion Answering

JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks

2026-02-06 · Lanbo Lin, Jiayao Liu, Tianyuan Yang, Li Cai 외 arxiv

Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide rigorous, reproducible assessment but fail to accommodate diverse valid response st…

PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning

2025-11-14 · Afra Feyza Akyürek, Advait Gosai, Chen Bo Calvin Zhang, Vipul Gupta 외 arxiv

Frontier model progress is often measured by academic benchmarks, which offer a limited view of performance in real-world professional contexts. Existing evaluations often fail to assess open-ended, economically conseque…

Xpertbench: Expert Level Tasks with Rubrics-Based Evaluation

2026-03-27 · Xue Liu, Xin Ma, Yuxin Ma, Yongchang Peng 외 arxiv

As Large Language Models (LLMs) exhibit plateauing performance on conventional benchmarks, a pivotal challenge persists: evaluating their proficiency in complex, open-ended tasks characterizing genuine expert-level cogni…