paper-with-me

홈 › Papers

DrugClaw and DrugAudit: A Primary-Source-Grounded Agent and Authority-Aware Benchmark for Drug-Information Question Answering

2026-05-31 · Qing Wang, Bo Li, Jialu Liang, Daling Shi, Bob Zhang, Qianqian Song arxiv

Drug-information question answering is a high-stakes setting where hallucinated facts can mislead clinical decision-making and the provenance of each cited fact matters as much as the fact itself. We present DrugClaw, a multi-agent retrieval-augmented system that queries a registry of drug and pharmacovigilance skills via a reflection-driven state-machine workflow and returns answers grounded in primary regulatory or peer-reviewed records. We also contribute DrugAudit, a 3,772-item authority-aware benchmark with an evaluation panel that scores upstream-of-gold source match, token-level semantic snippet overlap, and citation faithfulness under a dual-judge LLM-as-judge protocol with inter-judge kappa = 0.88 (almost-perfect). Across DrugAudit plus drug-related subsets of MedQA (751) and PubMedQA (512), DrugClaw is top-1 on every column of the headline table: composite Evidence Index under both judges, judge-mediated answer correctness, primary-source rate (0.918, +10.1 pp over next-best), faithfulness (0.887, +5.9 pp), MedQA (0.920), and PubMedQA (0.693).

📄 PDF Abstract BibTeX arXiv:2606.01434

Code (0)

등록된 구현이 없습니다.

Tasks

Question Answering

Similar Papers 제목 키워드 기반

Extending AI for Research to the Humanities: A Multi-Agent Framework for Evidence-Grounded Scholarship

2026-05-29 · Yating Pan, Jiajun Zhang, Jun Wang, Qi Su arxiv

LLM-based research agents have advanced rapidly in science and engineering, where research is organized around executable experiments, code, and quantitative signals. Humanities scholarship, however, requires interpretiv…

BIOGEN: Evidence-Grounded Multi-Agent Reasoning Framework for Transcriptomic Interpretation in Antimicrobial Resistance

2025-10-17 · Elias Hossain, Mehrdad Shoeibi, Ivan Garibay, Niloofar Yousefi arxiv

Interpreting gene clusters from RNA sequencing (RNA-seq) remains challenging, especially in antimicrobial resistance studies where mechanistic insight is important for hypothesis generation. Existing pathway enrichment m…

PhysicianBench: Evaluating LLM Agents in Real-World EHR Environments

2026-05-04 · Ruoqi Liu, Imran Q. Mohiuddin, Austin J. Schoeffler, Kavita Renduchintala 외 arxiv

We introduce PhysicianBench, a benchmark for evaluating LLM agents on physician tasks grounded in real clinical setting within electronic health record (EHR) environments. Existing medical agent benchmarks primarily focu…

Cite Before You Speak: Enhancing Context-Response Grounding in E-commerce Conversational LLM-Agents

2025-03-05 · Jingying Zeng, Hui Liu, Zhenwei Dai, Xianfeng Tang 외

With the advancement of conversational large language models (LLMs), several LLM-based Conversational Shopping Agents (CSA) have been developed to help customers smooth their online shopping. The primary objective in bui…

AttributeIn-Context LearningMisinformation

Less Interaction But More Explanation: A Communication Perspective on Agentic AI Interfaces

2026-05-02 · Eunchae Jang, S. Shyam Sundar arxiv

AI systems have long been expected to interact with users, answering questions, generating content, and continuing (social) conversations. Agentic AI, however, breaks from this expectation, as its primary objective is wo…