paper-with-me

Papers

MTBBench: A Multimodal Sequential Clinical Decision-Making Benchmark in Oncology

2025-11-25 · Kiril Vasilev, Alexandre Misrahi, Eeshaan Jain, Phil F Cheng, Petros Liakopoulos, Olivier Michielin, Michael Moor, Charlotte Bunne arxiv

Multimodal Large Language Models (LLMs) hold promise for biomedical reasoning, but current benchmarks fail to capture the complexity of real-world clinical workflows. Existing evaluations primarily assess unimodal, decontextualized question-answering, overlooking multi-agent decision-making environments such as Molecular Tumor Boards (MTBs). MTBs bring together diverse experts in oncology, where diagnostic and prognostic tasks require integrating heterogeneous data and evolving insights over time. Current benchmarks lack this longitudinal and multimodal complexity. We introduce MTBBench, an agentic benchmark simulating MTB-style decision-making through clinically challenging, multimodal, and longitudinal oncology questions. Ground truth annotations are validated by clinicians via a co-developed app, ensuring clinical relevance. We benchmark multiple open and closed-source LLMs and show that, even at scale, they lack reliability -- frequently hallucinating, struggling with reasoning from time-resolved data, and failing to reconcile conflicting evidence or different modalities. To address these limitations, MTBBench goes beyond benchmarking by providing an agentic framework with foundation model-based tools that enhance multi-modal and longitudinal reasoning, leading to task-level performance gains of up to 9.0% and 11.2%, respectively. Overall, MTBBench offers a challenging and realistic testbed for advancing multimodal LLM reasoning, reliability, and tool-use with a focus on MTB environments in precision oncology.

📄 PDF Abstract BibTeX arXiv:2511.20490

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments

2024-05-13 · Samuel Schmidgall, Rojin Ziaei, Carl Harris, Eduardo Reis 외

Evaluating large language models (LLM) in clinical scenarios is crucial to assessing their potential clinical utility. Existing benchmarks rely heavily on static question-answering, which does not accurately depict the c…

Decision MakingDiagnosticMedQAQuestion Answering+1

CURENet: Combining Unified Representations for Efficient Chronic Disease Prediction

2025-11-14 · Cong-Tinh Dao, Nguyen Minh Thao Phan, Jun-En Ding, Chenwei Wu 외 arxiv

Electronic health records (EHRs) are designed to synthesize diverse data types, including unstructured clinical notes, structured lab tests, and time-series visit data. Physicians draw on these multimodal and temporal so…

Correlational Dueling Bandits with Application to Clinical Treatment in Large Decision Spaces

2017-07-08 · Yanan Sui, Yisong Yue, Joel W. Burdick

We consider sequential decision making under uncertainty, where the goal is to optimize over a large decision space using noisy comparative feedback. This problem can be formulated as a $K$-armed Dueling Bandits problem …

Decision MakingDecision Making Under UncertaintySequential Decision Making

SAGEAgent: A Self-Evolving Agent for Cost-Aware Modality Acquisition in Multimodal Survival Prediction

2026-07-10 · Chongyu Qu, Can Cui, Zhengyi Lu, Junchao Zhu 외 arxiv

Does every cancer patient truly need a complete diagnostic workup for accurate survival prediction? In multimodal clinical oncology, diagnostic modalities follow a clinically mandated order of escalating burden -- from d…

Unified Models of Human Behavioral Agents in Bandits, Contextual Bandits and RL

2020-05-10 · Baihan Lin, Guillermo Cecchi, Djallel Bouneffouf, Jenna Reinen 외

Artificial behavioral agents are often evaluated based on their consistent behaviors and performance to take sequential actions in an environment to maximize some notion of cumulative reward. However, human decision maki…

Decision MakingLifelong learningMulti-Armed BanditsReinforcement Learning (RL)+1