paper-with-me

Papers

QuarkMedBench: A Real-World Scenario Driven Benchmark for Evaluating Large Language Models

2026-03-14 · Yao Wu, Kangping Yin, Liang Dong, Zhenxin Ma, Shuting Xu, Xuehai Wang, Yuxuan Jiang, Tingting Yu, Yunqing Hong, Jiayi Liu, Rianzhe Huang, Shuxin Zhao, Haiping Hu, Wen Shang, Jian Xu, Guanjun Jiang arxiv

While Large Language Models (LLMs) excel on standardized medical exams, high scores often fail to translate to high-quality responses for real-world medical queries. Current evaluations rely heavily on multiple-choice questions, failing to capture the unstructured, ambiguous, and long-tail complexities inherent in genuine user inquiries. To bridge this gap, we introduce QuarkMedBench, an ecologically valid benchmark tailored for real-world medical LLM assessment. We compiled a massive dataset spanning Clinical Care, Wellness Health, and Professional Inquiry, comprising 20,821 single-turn queries and 3,853 multi-turn sessions. To objectively evaluate open-ended answers, we propose an automated scoring framework that integrates multi-model consensus with evidence-based retrieval to dynamically generate 220,617 fine-grained scoring rubrics (~9.8 per query). During evaluation, hierarchical weighting and safety constraints structurally quantify medical accuracy, key-point coverage, and risk interception, effectively mitigating the high costs and subjectivity of human grading. Experimental results demonstrate that the generated rubrics achieve a 91.8% concordance rate with clinical expert blind audits, establishing highly dependable medical reliability. Crucially, baseline evaluations on this benchmark reveal significant performance disparities among state-of-the-art models when navigating real-world clinical nuances, highlighting the limitations of conventional exam-based metrics. Ultimately, QuarkMedBench establishes a rigorous, reproducible yardstick for measuring LLM performance on complex health issues, while its framework inherently supports dynamic knowledge updates to prevent benchmark obsolescence.

📄 PDF Abstract BibTeX arXiv:2603.13691

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Conditional Flow-VAE for Safety-Critical Traffic Scenario Generation

2026-05-06 · Zimu Gong, Brian Zhaoning Zhang, Chris Zhang, Kelvin Wong 외 arxiv

Safety-critical scenarios are essential for the development of autonomous vehicles (AVs) but are rare in real-world driving data. While simulation offers a way to generate such scenarios, manually designed test cases lac…

Autonomous Vehicles

ESF-Bench: Benchmarking Challenging Slot-Filling Scenarios for Real-World Enterprise Applications

2026-07-25 · Toby Liang, Gopal Sarda, Sagar Davasam, Vikas Yadav arxiv

The rapid rise of large language models (LLMs) has driven transformative adoption across enterprises. However, deploying these models in real-world settings presents unique challenges due to complex system constraints an…

Slot Filling

Systematic Benchmarking of SUMO Against Data-Driven Traffic Simulators

2025-12-20 · Erdao Liang arxiv

This paper presents a systematic benchmarking of the model-based microscopic traffic simulator SUMO against state-of-the-art data-driven traffic simulators using large-scale real-world datasets. Using the Waymo Open Moti…

Autonomous Driving

SenseJudge: Human-Centric Preference-Driven Judgment Framework

2026-06-02 · Rui Li, Junfeng Liu, Xiangwen Kong, Linhai Xu 외 arxiv

Large Language Models (LLMs) as judges across various scenarios such as assessing model responses is becoming an increasingly accepted paradigm. However, existing judgment approaches often rely on trained judgers using f…

BO4Mob: Bayesian Optimization Benchmarks for High-Dimensional Urban Mobility Problem

2025-10-21 · Seunghee Ryu, Donghoon Kwon, Seongjin Choi, Aryan Deshwal 외 arxiv

We introduce \textbf{BO4Mob}, a new benchmark framework for high-dimensional Bayesian Optimization (BO), driven by the challenge of origin-destination (OD) travel demand estimation in large urban road networks. Estimatin…