paper-with-me

Papers

Uncovering Neural Scaling Laws in Molecular Representation Learning

2023-09-15 · NeurIPS 2023 11 · Dingshuo Chen, Yanqiao Zhu, Jieyu Zhang, Yuanqi Du, ZHIXUN LI, Qiang Liu, Shu Wu, Liang Wang

Molecular Representation Learning (MRL) has emerged as a powerful tool for drug and materials discovery in a variety of tasks such as virtual screening and inverse design. While there has been a surge of interest in advancing model-centric techniques, the influence of both data quantity and quality on molecular representations is not yet clearly understood within this field. In this paper, we delve into the neural scaling behaviors of MRL from a data-centric viewpoint, examining four key dimensions: (1) data modalities, (2) dataset splitting, (3) the role of pre-training, and (4) model capacity. Our empirical studies confirm a consistent power-law relationship between data volume and MRL performance across these dimensions. Additionally, through detailed analysis, we identify potential avenues for improving learning efficiency. To challenge these scaling laws, we adapt seven popular data pruning strategies to molecular data and benchmark their performance. Our findings underline the importance of data-centric MRL and highlight possible directions for future research.

📄 PDF Abstract BibTeX arXiv:2309.15123

Code (2)

Data-reindeer/MolScaling 공식 구현 pytorch
data-reindeer/nsl_mrl 공식 구현 pytorch

Tasks

molecular representationRepresentation Learning

Methods 이 논문이 사용한 방법론

Pruning 설명 없음

Similar Papers 제목 키워드 기반

Unveiling Scaling Behaviors in Molecular Language Models: Effects of Model Size, Data, and Representation

2026-01-30 · Dong Xu, Qihua Pan, Sisi Yuan, Jianqiang Li 외 arxiv

Molecular generative models, often employing GPT-style language modeling on molecular string representations, have shown promising capabilities when scaled to large datasets and model sizes. However, it remains unclear a…

Uncovering Scaling Laws for Large Language Models via Inverse Problems

2025-09-09 · Arun Verma, Zhaoxuan Wu, Zijian Zhou, Xiaoqiang Lin 외 arxiv

Large Language Models (LLMs) are large-scale pretrained models that have achieved remarkable success across diverse domains. These successes have been driven by unprecedented complexity and scale in both data and computa…

Measurement noise scaling laws for cellular representation learning

2025-03-04 · Gokul Gowri, Peng Yin, Allon M. Klein

Deep learning scaling laws predict how performance improves with increased model and dataset size. Here we identify measurement noise in data as another performance scaling axis, governed by a distinct logarithmic law. W…

image-classificationImage ClassificationRepresentation Learning

Universal scaling laws in quantum-probabilistic machine learning by tensor network towards interpreting representation and generalization powers

2024-10-13 · Sheng-Chen Bai, Shi-Ju Ran

Interpreting the representation and generalization powers has been a long-standing issue in the field of machine learning (ML) and artificial intelligence. This work contributes to uncovering the emergence of universal s…

Omni-Mol: Exploring Universal Convergent Space for Omni-Molecular Tasks

2025-02-03 · Chengxin Hu, Hao Li, Yihe Yuan, Zezheng Song 외

Building generalist models has recently demonstrated remarkable capabilities in diverse scientific domains. Within the realm of molecular learning, several studies have explored unifying diverse tasks across diverse doma…

Active Learning