paper-with-me

홈 › Papers

Simulation as Supervision: Mechanistic Pretraining for Scientific Discovery

2025-07-11 · Carson Dudley, Reiden Magdaleno, Christopher Harding, Marisa Eisenberg arxiv

Scientific modeling faces a tradeoff between the interpretability of mechanistic theory and the predictive power of machine learning. While existing hybrid approaches have made progress by incorporating domain knowledge into machine learning methods as functional constraints, they can be limited by a reliance on precise mathematical specifications. When the underlying equations are partially unknown or misspecified, enforcing rigid constraints can introduce bias and hinder a model's ability to learn from data. We introduce Simulation-Grounded Neural Networks (SGNNs), a framework that incorporates scientific theory by using mechanistic simulations as training data for neural networks. By pretraining on diverse synthetic corpora that span multiple model structures and realistic observational noise, SGNNs internalize the underlying dynamics of a system as a structural prior. We evaluated SGNNs across multiple disciplines, including epidemiology, ecology, social science, and chemistry. In forecasting tasks, SGNNs outperformed both standard data-driven baselines and physics-constrained hybrid models. They nearly tripled the forecasting skill of the average CDC models in COVID-19 mortality forecasts and accurately forecasted high-dimensional ecological systems. SGNNs demonstrated robustness to model misspecification, performing well even when trained on data with incorrect assumptions. Our framework also introduces back-to-simulation attribution, a method for mechanistic interpretability that explains real-world dynamics by identifying their most similar counterparts within the simulated corpus. By unifying these techniques into a single framework, we demonstrate that diverse mechanistic simulations can serve as effective training data for robust scientific inference.

📄 PDF Abstract BibTeX arXiv:2507.08977

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Autonomous Mechanistic Reasoning in Virtual Cells

2026-04-13 · Yunhui Jang, Lu Zhu, Jake Fawkes, Alisandra Kaye Denton 외 arxiv

Large language models (LLMs) have recently gained significant attention as a promising approach to accelerate scientific discovery. However, their application in open-ended scientific domains such as biology remains limi…

Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence

2026-08-12 · Mengru Wang, Junfeng Fang, Shuofei Qiao, Zhenqian Xu 외 hf

AI models have achieved remarkable success across diverse domains, yet the mechanisms underlying their capabilities and the risks they may pose remain poorly understood. As AI development becomes faster and increasingly …

From Observation to Insight: Mechanistic World Models and the Quest for Autonomous Discovery

2026-07-14 · Ingmar Posner, Anson Lei, Bernhard Schölkopf arxiv

Recent advances in foundation models have transformed AI for Science, enabling remarkably accurate predictive performance across domains ranging from protein folding to weather forecasting. Yet prediction alone does not …

Representation LearningWeather Forecasting

Discovering Mechanistic Models of Neural Activity: System Identification in an in Silico Zebrafish

2026-02-04 · Jan-Matthis Lueckmann, Viren Jain, Michał Januszewski arxiv

Constructing mechanistic models of neural circuits is a fundamental goal of neuroscience, yet verifying such models is limited by the lack of ground truth. To rigorously test model discovery, we establish an in silico te…

Position: Prioritize Identifying Structure, Not Complex Models, for Scientific Discovery

2026-05-30 · Tyler H. McCormick arxiv

Modern Machine Learning (ML) and Artificial Intelligence (AI) models, especially large language models (LLMs), are increasingly used to generate scientific hypotheses and mechanistic explanations from observational data.…