paper-with-me

홈 › Papers

FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment

2026-03-20 · Betty Xiong, Jillian Fisher, Benjamin Newman, Meng Hu, Shivangi Gupta, Yejin Choi, Lanyan Fang, Russ B Altman arxiv

We introduce an expert curated, real-world benchmark for evaluating document-grounded question-answering (QA) motivated by generic drug assessment, using the U.S. Food and Drug Administration (FDA) drug label documents. Drug labels contain rich but heterogeneous clinical and regulatory information, making accurate question answering difficult for current language models. In collaboration with FDA regulatory assessors, we introduce FDARxBench, and construct a multi-stage pipeline for generating high-quality, expert curated, QA examples spanning factual, multi-hop, and refusal tasks, and design evaluation protocols to assess both open-book and closed-book reasoning. Experiments across proprietary and open-weight models reveal substantial gaps in factual grounding, long-context retrieval, and safe refusal behavior. While motivated by FDA generic drug assessment needs, this benchmark also provides a substantial foundation for challenging regulatory-grade evaluation of label comprehension. The benchmark is designed to support evaluation of LLM behavior on drug-label questions.

📄 PDF Abstract BibTeX arXiv:2603.19539

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

MedBrowseComp: Benchmarking Medical Deep Research and Computer Use

2025-05-20 · Shan Chen, Pedro Moreira, Yuxin Xiao, Sam Schmidgall 외

Large language models (LLMs) are increasingly envisioned as decision-support tools in clinical practice, yet safe clinical reasoning demands integrating heterogeneous knowledge bases -- trials, primary studies, regulator…

Benchmarking

EURO-5K: When Does Domain Pretraining Matter? Benchmarking Transformers for EU Reporting Obligation Extraction

2026-06-02 · Marios Koniaris, Vasileios Kotronis, Eugenia Giannini, Panayiotis Tsanakas arxiv

Extracting reporting obligations from EU legislation is critical for assessing and reducing regulatory reporting burden. However, distinguishing reporting requirements from structurally similar provisions requires specia…

Domain Adaptation

MedBench v4: A Robust and Scalable Benchmark for Evaluating Chinese Medical Language Models, Multimodal Models, and Intelligent Agents

2025-11-18 · Jinru Ding, Lu Lu, Chao Ding, Mouxiao Bian 외 arxiv

Recent advances in medical large language models (LLMs), multimodal models, and agents demand evaluation frameworks that reflect real clinical workflows and safety constraints. We present MedBench v4, a nationwide, cloud…

Multimodal Reasoning

Approaches for benchmarking single-cell gene regulatory network inference methods

2023-07-17 · Yasin Uzun

Gene regulatory networks are powerful tools for modeling interactions among genes to regulate their expression for homeostasis and differentiation. Single-cell sequencing offers a unique opportunity to build these networ…

Benchmarking

PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning

2026-09-04 · Taegyun Kim, Youngwook Ham, Jungwook Rhim, Ju-Hyun An 외 arxiv

We introduce PetQA, a Korean long-form question-answering (QA) benchmark for evaluating veterinary knowledge and clinical reasoning in large language models (LLMs) and large vision-language models (LVLMs). PetQA contains…