paper-with-me

Papers

Towards a Deterministic Math Solver for Clinical Language Models

2026-09-09 · Felipe Ocampo Osorio, Sebastián Andrés Cajas Ordoñez, Maximin Lange, Rafi Al Attrach, Sahil Kapadia, Zakaria Laouabdia Sellami, Angelo Antonio Talio, Leo Anthony Celi hf

Large language models are unreliable at arithmetic, which is a problem for clinical calculators where a single numerical error changes the recommendation. The standard response is to hardcode each calculator as a validated function, one at a time. We test an alternative: the model does not calculate. Instead, it writes case-specific Python that a restricted local executor runs as a deterministic solver, and the model's task reduces to deciding how to use it. We evaluate this Program-Solve interface on MedCalc-Bench Verified (1,100 cases, 55 calculators) against direct model arithmetic and a hand-written 22-calculator library, using Qwen2.5-7B and Qwen2.5-32B-AWQ, after auditing the benchmark's formulas against current clinical guidelines and flagging 16 of 55 with version, use or coefficient concerns. With formulas and gold variables supplied and both routes reading the whole note, handing off to the solver is not a reliable advantage at 7B (75.31% against 72.02%, a paired +3.29 points with a 95% calculator-cluster interval of [-3.49, 10.38]) but is one at 32B (90.53% against 83.47%, +7.05 [0.47, 14.60], clear of zero). The hand-written library is exact on its 440 supported cases but abstains elsewhere (40.0% overall). Adding an executor thus helps some open-weight models more than others even under matched formula, variable and note access, and is not a substitute for verified formulas or reliable variable extraction either way.

📄 PDF Abstract BibTeX arXiv:2609.10728

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Toward Guarantees for Clinical Reasoning in Vision Language Models via Formal Verification

2026-02-27 · Vikash Singh, Debargha Ganguly, Haotian Yu, Chengwei Zhou 외 arxiv

Vision-language models (VLMs) show promise in drafting radiology reports, yet they frequently suffer from logical inconsistencies, generating diagnostic impressions unsupported by their own perceptual findings or missing…

Clinical Knowledge

SPL: Orchestrating Workflows with Declarative Deterministic-Probabilistic Composition

2026-07-06 · Wen G. Gong arxiv

We present SPL (Structured Prompt Language), a declarative language that composes deterministic and probabilistic computation modes in a single specification. While existing frameworks separate these -- orchestration sys…

Latent Generative Solvers for Generalizable Long-Term Physics Simulation

2026-02-11 · Zituo Chen, Sili Deng arxiv

Reliable physics simulation demands two capabilities that today's neural PDE solvers do not deliver together: generalization across heterogeneous PDE families, and stability under long autoregressive rollouts. Determinis…

Structure-Unified M-Tree Coding Solver for MathWord Problem

2022-10-22 · Bin Wang, Jiangzhou Ju, Yang Fan, Xinyu Dai 외

As one of the challenging NLP tasks, designing math word problem (MWP) solvers has attracted increasing research attention for the past few years. In previous work, models designed by taking into account the properties o…

Math

Measure Anatomical Thickness from Cardiac MRI with Deep Neural Networks

2020-08-25 · Qiaoying Huang, Eric Z. Chen, Hanchao Yu, Yimo Guo 외

Accurate estimation of shape thickness from medical images is crucial in clinical applications. For example, the thickness of myocardium is one of the key to cardiac disease diagnosis. While mathematical models are avail…