paper-with-me

Papers

Benchmarking Clinical Decision Support Search

2018-01-29 · Nguyen Vincent, Karimi Sarvnaz, Falamaki Sara, Paris Cecile

Finding relevant literature underpins the practice of evidence-based medicine. From 2014 to 2016, TREC conducted a clinical decision support track, wherein participants were tasked with finding articles relevant to clinical questions posed by physicians. In total, 87 teams have participated over the past three years, generating 395 runs. During this period, each team has trialled a variety of methods. While there was significant overlap in the methods employed by different teams, the results were varied. Due to the diversity of the platforms used, the results arising from the different techniques are not directly comparable, reducing the ability to build on previous work. By using a stable platform, we have been able to compare different document and query processing techniques, allowing us to experiment with different search parameters. We have used our system to reproduce leading teams runs, and compare the results obtained. By benchmarking our indexing and search techniques, we can statistically test a variety of hypotheses, paving the way for further research.

📄 PDF Abstract BibTeX arXiv:1801.09322

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesBenchmarkingDiversity

Similar Papers 제목 키워드 기반

MedCase-Structured: A Text-to-FHIR Dataset for Benchmarking Diagnostic Reasoning in Clinically Realistic EHR Settings

2026-05-28 · Valentina Bui Muti, Eugénie Dulout, Ziquan Fu arxiv

Large language models (LLMs) show promise for clinical reasoning and decision support, but evaluation in structured, electronic health record-congruent settings remains limited. Existing benchmarks often rely on static d…

MedMASLab: A Unified Orchestration Framework for Benchmarking Multimodal Medical Multi-Agent Systems

2026-03-10 · Yunhang Qian, Xiaobin Hu, Jiaquan Yu, Siyang Xin 외 arxiv

While Multi-Agent Systems (MAS) show potential for complex clinical decision support, the field remains hindered by architectural fragmentation and the lack of standardized multimodal integration. Current medical MAS res…

Visual Grounding

PsychBench: A comprehensive and professional benchmark for evaluating the performance of LLM-assisted psychiatric clinical practice

2025-02-28 · Shuyu Liu, Ruoxi Wang, Ling Zhang, Xuequan Zhu 외

The advent of Large Language Models (LLMs) offers potential solutions to address problems such as shortage of medical resources and low diagnostic consistency in psychiatric clinical practice. Despite this potential, a r…

BenchmarkingDiagnostic

Yesil o1 Pro: Evidence-Based AI Model for Health and Benchmarking in Clinical Decision Support

2025-02-15 · JMIR Preprints 2025 2 · Yusuf Yesil

Background: Integrating evidence-based approaches in healthcare and artificial intelligence (AI) is crucial for enhancing clinical decision-making and patient safety. Yesil o1 Pro is a specialized large language model (…

BenchmarkingEpidemiologyLarge Language Model

Benchmarking Early Deterioration Prediction Across Hospital-Rich and MCI-Like Emergency Triage Under Constrained Sensing

2026-02-09 · KMA Solaiman, Joshua Sebastian, Karma Tobden arxiv

Emergency triage decisions are made under severe information constraints, yet most data-driven deterioration models are evaluated using signals unavailable during initial assessment. We present a leakage-aware benchmarki…