paper-with-me

홈 › Papers

From Stochastic Answers to Verifiable Reasoning: Interpretable Decision-Making with LLM-Generated Code

2026-02-28 · Anirudh Jaidev Mahesh, Ben Griffin, Fuat Alican, Joseph Ternasky, Zakari Salifu, Kelvin Amoaba, Yagiz Ihlamur, Aaron Ontoyin Yin, Aikins Laryea, Afriyie Samuel, Yigit Ihlamur arxiv

Large language models (LLMs) are increasingly used for high-stakes decision-making, yet existing approaches struggle to reconcile scalability, interpretability, and reproducibility. Black-box models obscure their reasoning, while recent LLM-based rule systems rely on per-sample evaluation, causing costs to scale with dataset size and introducing stochastic, hallucination-prone outputs. We propose reframing LLMs as code generators rather than per-instance evaluators. A single LLM call generates executable, human-readable decision logic that runs deterministically over structured data, eliminating per-sample LLM queries while enabling reproducible and auditable predictions. We combine code generation with automated statistical validation using precision lift, binomial significance testing, and coverage filtering, and apply cluster-based gap analysis to iteratively refine decision logic without human annotation. We instantiate this framework in venture capital founder screening, a rare-event prediction task with strong interpretability requirements. On VCBench, a benchmark of 4,500 founders with a 9% base success rate, our approach achieves 37.5% precision and an F0.5 score of 25.0%, outperforming GPT-4o (at 30.0% precision and an F0.5 score of 25.7%) while maintaining full interpretability. Each prediction traces to executable rules over human-readable attributes, demonstrating verifiable and interpretable LLM-based decision-making in practice.

📄 PDF Abstract BibTeX arXiv:2603.13287

Code (0)

등록된 구현이 없습니다.

Tasks

Code Generation

Similar Papers 제목 키워드 기반

Trade-R1: Bridging Verifiable Rewards to Stochastic Environments via Process-Level Reasoning Verification

2026-01-07 · Rui Sun, Yifan Sun, Sheng Xu, Li Zhao 외 arxiv

Reinforcement Learning (RL) has enabled Large Language Models (LLMs) to achieve remarkable reasoning in domains like mathematics and coding, where verifiable rewards provide clear signals. However, extending this paradig…

Reinforcement Learning

Bradley-Terry Policy Optimization for Generative Preference Modeling

2025-10-17 · Shengyu Feng, Yun He, Shuang Ma, Beibin Li 외 arxiv

Reinforcement learning (RL) has recently proven effective at scaling chain-of-thought (CoT) reasoning in large language models for tasks with verifiable answers. However, extending RL-based thought training to more gener…

Reinforcement Learning

CoReTab: Improving Multimodal Table Understanding with Code-driven Reasoning

2026-01-27 · Van-Quang Nguyen, Takayuki Okatani arxiv

Existing datasets for multimodal table understanding, such as MMTab, primarily provide short factual answers without explicit multi-step reasoning supervision. Models trained on these datasets often generate brief respon…

Question AnsweringFact Verification

Composing Neural Learning and Symbolic Reasoning with an Application to Visual Discrimination

2019-07-12 · Adithya Murali, Atharva Sehgal, Paul Krogmeier, P. Madhusudan

We consider the problem of combining machine learning models to perform higher-level cognitive tasks with clear specifications. We propose the novel problem of Visual Discrimination Puzzles (VDP) that requires finding in…

Few-Shot Learning

From Models to Microtheories: Distilling a Model's Topical Knowledge for Grounded Question Answering

2024-12-23 · Nathaniel Weir, Bhavana Dalvi Mishra, Orion Weller, Oyvind Tafjord 외

Recent reasoning methods (e.g., chain-of-thought, entailment reasoning) help users understand how language models (LMs) answer a single question, but they do little to reveal the LM's overall understanding, or "theory," …

Question Answering