paper-with-me

홈 › Papers

Automatic Construction of Clinical Scoring Systems with LLM Agents

2026-01-29 · Silas Ruhrberg Estévez, Christopher Chiu, Mihaela van der Schaar arxiv

Modern clinical practice relies on evidence-based guidelines implemented as compact scoring systems composed of a small number of interpretable decision rules. While machine-learning models achieve strong performance, many fail to translate into routine clinical use due to misalignment with workflow constraints such as memorability, auditability, and bedside execution. We argue that this gap arises not from insufficient predictive power, but from optimizing over model classes that are incompatible with guideline deployment. Deployable guidelines often take the form of unit-weighted clinical checklists, formed by thresholding the sum of binary rules, but learning such scores requires searching an exponentially large discrete space of possible rule sets. We introduce AgentScore, which performs semantically guided optimization in this space by using LLMs to propose candidate rules and a deterministic, data-grounded verification-and-selection loop to enforce statistical validity and deployability constraints. Across eight clinical prediction tasks, AgentScore outperforms existing score-generation methods and achieves AUROC comparable to more flexible interpretable models despite operating under stronger structural constraints. On two additional externally validated tasks, AgentScore achieves higher discrimination than established guideline-based scores.

📄 PDF Abstract BibTeX arXiv:2601.22324

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Automatic Sleep Staging of EEG Signals: Recent Development, Challenges, and Future Directions

2021-11-03 · Huy Phan, Kaare Mikkelsen

Modern deep learning holds a great potential to transform clinical practice on human sleep. Teaching a machine to carry out routine tasks would be a tremendous reduction in workload for clinicians. Sleep staging, a funda…

EEGElectroencephalogram (EEG)Sleep Staging

AgentsEval: Clinically Faithful Evaluation of Medical Imaging Reports via Multi-Agent Reasoning

2026-01-23 · Suzhong Fu, Jingqi Dong, Xuan Ding, Rui Sun 외 arxiv

Evaluating the clinical correctness and reasoning fidelity of automatically generated medical imaging reports remains a critical yet unresolved challenge. Existing evaluation methods often fail to capture the structured …

Medical Report Generation

Personalized 3D Myocardial Infarct Geometry Reconstruction from Cine MRI with Explicit Cardiac Motion Modeling

2025-07-21 · Yilin Lyu, Fan Yang, Xiaoyue Liu, Zichen Jiang 외 arxiv

Accurate representation of myocardial infarct geometry is crucial for patient-specific cardiac modeling in MI patients. While Late gadolinium enhancement (LGE) MRI is the clinical gold standard for infarct detection, it …

Automated Scoring of Clinical Expressive Language Evaluation Tasks

2020-07-01 · WS 2020 7 · Yiyi Wang, Emily Prud{'}hommeaux, Meysam Asgari, Jill Dolata

Many clinical assessment instruments used to diagnose language impairments in children include a task in which the subject must formulate a sentence to describe an image using a specific target word. Because producing se…

DiversityMachine TranslationSentenceTransfer Learning+2

GEMA-Score: Granular Explainable Multi-Agent Score for Radiology Report Evaluation

2025-03-07 · Zhenxuan Zhang, Kinhei Lee, Weihang Deng, Huichi Zhou 외

Automatic medical report generation supports clinical diagnosis, reduces the workload of radiologists, and holds the promise of improving diagnosis consistency. However, existing evaluation metrics primarily assess the a…

Large Language ModelMedical Report GenerationNER