paper-with-me

홈 › Papers

Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search

2026-08-16 · Zhongwei Yu, Yan Song, Xue Yan, Anjie Liu, Xingyu Lu, Yihang Chen, Huichi Zhou, Siyuan Guo, Luoyang Sun, Sihan Chen, Xiangning Yu, Jun Wang hf

Scientific discovery often involves optimising expensive-to-evaluate objectives over vast, structured, and open-ended hypothesis spaces, such as molecules, protein sequences, and computer programs. Generative models such as large language models (LLMs) provide expressive priors over such spaces, but their likelihoods and self-assessments are unreliable proxies for the objectives and calibrated epistemic uncertainty, especially for novel candidates outside the observed data distribution. We introduce the Large Discovery Model (LDM), an empirically grounded recurrent architecture that couples a generative model with a Bayesian non-parametric reward surrogate model. The generative model proposes and refines candidate designs, while the surrogate predicts their performance and quantifies uncertainty, yielding an uncertainty-aware value that guides candidate generation, refinement, and selection. The discovery memory and the surrogate model are continually updated as each new experimental observation arrives. We evaluate LDM on three scenarios spanning different design modalities and objectives, including neural-network training, antibody design, and molecular optimisation. Compared to LLM-only reflection or traditional statistical search across these domains, LDM achieves a 2.4times greater reduction in validation BPB, an 18.2% relative decrease in binding energy, and more than 60% relative gains in molecular multi-objective performance. These results suggests that LDM could serve as a general-purpose discovery engine for effective search over open-ended hypothesis spaces.

📄 PDF Abstract BibTeX arXiv:2608.15669

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Towards Autonomous Mechanistic Reasoning in Virtual Cells

2026-04-13 · Yunhui Jang, Lu Zhu, Jake Fawkes, Alisandra Kaye Denton 외 arxiv

Large language models (LLMs) have recently gained significant attention as a promising approach to accelerate scientific discovery. However, their application in open-ended scientific domains such as biology remains limi…

Discovery Foundation Models: Toward Open-Ended Discovery Intelligence

2026-09-14 · Ling Yang, Zhenfei Yin, Yingcheng Wu hf

Foundation models have progressed from learning and reasoning over existing knowledge, to increasingly learning through action, tool use, and outcome feedback. We argue that the next frontier is a further transition: fro…

AI CFD Scientist: Toward Open-Ended Computational Fluid Dynamics Discovery with Physics-Aware AI Agents

2026-05-07 · Nithin Somasekharan, Rabi Pathak, Manushri Dhanakoti, Tingwen Zhang 외 arxiv

Recent LLM-based agents have closed substantial portions of the scientific discovery loop in software-only machine-learning research, in chemistry, and in biology. Extending the same loop to high-fidelity physical simula…

Towards Open-Ended Discovery for Low-Resource NLP

2025-09-22 · Bonaventure F. P. Dossou, Henri Aïdasso arxiv

Natural Language Processing (NLP) for low-resource languages remains fundamentally constrained by the lack of textual corpora, standardized orthographies, and scalable annotation pipelines. While recent advances in large…

Cross-Lingual Transfer

LiFT: Unsupervised Reinforcement Learning with Foundation Models as Teachers

2023-12-14 · Taewook Nam, Juyong Lee, Jesse Zhang, Sung Ju Hwang 외

We propose a framework that leverages foundation models as teachers, guiding a reinforcement learning agent to acquire semantically meaningful behavior without human feedback. In our framework, the agent receives task in…

Language ModelingLanguage Modellingreinforcement-learningReinforcement Learning+1