paper-with-me

홈 › Papers

Position: Prioritize Identifying Structure, Not Complex Models, for Scientific Discovery

2026-05-30 · Tyler H. McCormick arxiv

Modern Machine Learning (ML) and Artificial Intelligence (AI) models, especially large language models (LLMs), are increasingly used to generate scientific hypotheses and mechanistic explanations from observational data. This position paper argues that in the high-dimensional proxy regimes where modern ML excels, mechanistic learning is generically underdetermined: many incompatible mechanisms induce essentially the same observational relationships on the support of the data, so predictive success and coherent explanations are insufficient evidence of mechanism discovery. This underdetermination becomes uniquely hazardous with large language models (LLMs), which tend to collapse large equivalence classes of explanations into a single fluent narrative. This paper proposes concrete standards for ``mechanistic ML,'' and argues these norms are necessary if LLM-centered workflows are to support science rather than merely simulate it.

📄 PDF Abstract BibTeX arXiv:2606.02632

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Complex Networks and the Drug Repositioning Problem

2026-02-26 · Felipe Bivort Haiek arxiv

In this Master's thesis, the graph properties of a multi-level drug-protein network are studied, as well as how the network's shape has informed discoveries over the years, identifying primarily crawling discoveries and …

Enabling knowledge discovery in natural hazard engineering datasets on DesignSafe

2023-04-21 · Chahak Mehta, Krishna Kumar

Data-driven discoveries require identifying relevant data relationships from a sea of complex, unstructured, and heterogeneous scientific data. We propose a hybrid methodology that extracts metadata and leverages scienti…

Knowledge Graphs

Position: Ideas Should be the Center of Machine Learning Research

2026-05-14 · Jairo Diaz-Rodriguez arxiv

Machine learning research increasingly bifurcates into two disconnected modes: benchmark-driven engineering that prioritizes metrics over understanding, and idealized theory that often fails to transfer to modern systems…

CiteQA@CLSciSumm 2020

2020-11-01 · EMNLP (sdp) 2020 11 · Anjana Umapathy, Karthik Radhakrishnan, Kinjal Jain, Rahul Singh

In academic publications, citations are used to build context for a concept by highlighting relevant aspects from reference papers. Automatically identifying referenced snippets can help researchers swiftly isolate princ…

Articles

Epistemic Uncertainty for Test-Time Discovery

2026-05-11 · Kainat Riaz, Muhammad Ahmed Mohsin, Ahsan Bilal, Muhammad Umer 외 arxiv

Automated scientific discovery using large language models relies on identifying genuinely novel solutions. Standard reinforcement learning penalizes high-variance mutations, which leads the policy to prioritize familiar…

Reinforcement Learning