paper-with-me

홈 › Papers

Beyond Perfect Scores: Proof-by-Contradiction for Trustworthy Machine Learning

2026-01-10 · Dushan N. Wadduwage, Dineth Jayakody, Leonidas Zimianitis arxiv

Machine learning (ML) models show strong promise for new biomedical prediction tasks, but concerns about trustworthiness have hindered their clinical adoption. In particular, it is often unclear whether a model relies on true clinical cues or on spurious hierarchical correlations in the data. This paper introduces a simple yet broadly applicable trustworthiness test grounded in stochastic proof-by-contradiction. Instead of just showing high test performance, our approach trains and tests on spurious labels carefully permuted based on a potential outcomes framework. A truly trustworthy model should fail under such label permutation; comparable accuracy across real and permuted labels indicates overfitting, shortcut learning, or data leakage. Our approach quantifies this behavior through interpretable Fisher-style p-values, which are well understood by domain experts across medical and life sciences. We evaluate our approach on multiple new bacterial diagnostics to separate tasks and models learning genuine causal relationships from those driven by dataset artifacts or statistical coincidences. Our work establishes a foundation to build rigor and trust between ML and life-science research communities, moving ML models one step closer to clinical adoption.

📄 PDF Abstract BibTeX arXiv:2601.06704

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Contradictions

2025-09-07 · Yang Xu, Shuwei Chen, Xiaomei Zhong, Jun Liu 외 arxiv

Trustworthy AI requires reasoning systems that are not only powerful but also transparent and reliable. Automated Theorem Proving (ATP) is central to formal reasoning, yet classical binary resolution remains limited, as …

Automated Theorem Proving

LegalWiz: A Multi-Agent Generation Framework for Contradiction Detection in Legal Documents

2025-10-03 · Ananya Mantravadi, Shivali Dalmia, Olga Pospelova, Abhishek Mukherji 외 arxiv

Retrieval-Augmented Generation (RAG) integrates large language models (LLMs) with external sources, but unresolved contradictions in retrieved evidence often lead to hallucinations and legally unsound outputs. Benchmarks…

Deduction Theorem: The Problematic Nature of Common Practice in Game Theory

2019-07-31 · Holger I. Meinhardt

We consider the Deduction Theorem used in the literature of game theory to run a purported proof by contradiction. In the context of game theory, it is stated that if we have a proof of $\phi \vdash \varphi$, then we als…

Faults in Our Formal Benchmarking: Dataset Defects and Evaluation Failures in Lean Theorem Proving

2026-06-28 · Pawan Sasanka Ammanamanchi, Siddharth Bhat, Stella Biderman arxiv

Benchmarks for LLM-assisted theorem proving in Lean are often treated as intrinsically reliable because every solved instance comes with a machine-checked proof. However, the kernel only checks that a proof establishes a…

Improving Bot Response Contradiction Detection via Utterance Rewriting

2022-07-25 · SIGDIAL (ACL) 2022 9 · Di Jin, Sijia Liu, Yang Liu, Dilek Hakkani-Tur

Though chatbots based on large neural models can often produce fluent responses in open domain conversations, one salient error type is contradiction or inconsistency with the preceding conversation turns. Previous work …

Natural Language Inference