paper-with-me

Papers

MedSkillAudit: A Domain-Specific Audit Framework for Medical Research Agent Skills

2026-04-22 · Yingyong Hou, Xinyuan Lao, Huimei Wang, Qianyu Yao, Wei Chen, Bocheng Huang, Fei Sun, Yuxian Lv, Weiqi Lei, Xueqian Wen, Pengfei Xia, Zhujun Tan, Shengyang Xie arxiv

Background: Agent skills are increasingly deployed as modular, reusable capability units in AI agent systems. Medical research agent skills require safeguards beyond general-purpose evaluation, including scientific integrity, methodological validity, reproducibility, and boundary safety. This study developed and preliminarily evaluated a domain-specific audit framework for medical research agent skills, with a focus on reliability against expert review. Methods: We developed MedSkillAudit (skill-auditor@1.0), a layered framework assessing skill release readiness before deployment. We evaluated 75 skills across five medical research categories (15 per category). Two experts independently assigned a quality score (0-100), an ordinal release disposition (Production Ready / Limited Release / Beta Only / Reject), and a high-risk failure flag. System-expert agreement was quantified using ICC(2,1) and linearly weighted Cohen's kappa, benchmarked against the human inter-rater baseline. Results: The mean consensus quality score was 72.4 (SD = 13.0); 57.3% of skills fell below the Limited Release threshold. MedSkillAudit achieved ICC(2,1) = 0.449 (95% CI: 0.250-0.610), exceeding the human inter-rater ICC of 0.300. System-consensus score divergence (SD = 9.5) was smaller than inter-expert divergence (SD = 12.4), with no directional bias (Wilcoxon p = 0.613). Protocol Design showed the strongest category-level agreement (ICC = 0.551); Academic Writing showed a negative ICC (-0.567), reflecting a structural rubric-expert mismatch. Conclusions: Domain-specific pre-deployment audit may provide a practical foundation for governing medical research agent skills, complementing general-purpose quality checks with structured audit workflows tailored to scientific use cases.

📄 PDF Abstract BibTeX arXiv:2604.20441

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

TRACE-Seg3D: Counterfactual Context Auditing For Robust 3D Glioma Segmentation Under Institutional Shift

2026-07-08 · Nguyen Linh Dan Le, Nguyen Pham Hoang Le, Tran Dang Khoi arxiv

Medical image segmentation models can achieve strong benchmark performance while remaining sensitive to scanner, protocol, and institutional variation. These context shifts alter image appearance without changing the und…

Medical Image Segmentation

A multimodal slice discovery framework for systematic failure detection and explanation in medical image classification

2026-02-27 · Yixuan Liu, Kanwal K. Bhatia, Ahmed E. Fetit arxiv

Despite advances in machine learning-based medical image classifiers, the safety and reliability of these systems remain major concerns in practical settings. Existing auditing approaches mainly rely on unimodal features…

Medical Image ClassificationExplanation Generation

Semi-structured LLM Reasoners Can Be Rigorously Audited

2025-05-30 · Jixuan Leng, Cassandra A. Cohen, Zhixian Zhang, Chenyan Xiong 외

As Large Language Models (LLMs) become increasingly capable at reasoning, the problem of "faithfulness" persists: LLM "reasoning traces" can contain errors and omissions that are difficult to detect, and may obscure bias…

Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs

2025-11-02 · Yan Shu, Chi Liu, Robin Chen, Derek Li 외 arxiv

Multimodal Large Language Models (MLLMs) have demonstrated remarkable effectiveness in various general-domain scenarios, such as visual question answering and image captioning. Recently, researchers have increasingly foc…

Visual Question AnsweringImage CaptioningVisual Reasoning

Beyond Endpoint Gains: A Weight-Delta Audit of Medical Specialization

2026-08-21 · Praphul Singh, Shanu Kumar, Akshat Agarwal arxiv

Specialist language models are usually understood through endpoint gains: the generalist scores lower, the specialist scores higher, and the difference is treated as evidence of specialization. This leaves the released u…