paper-with-me

홈 › Papers

On the Interplay of Subset Selection and Informed Graph Neural Networks

2023-06-15 · Niklas Breustedt, Paolo Climaco, Jochen Garcke, Jan Hamaekers, Gitta Kutyniok, Dirk A. Lorenz, Rick Oerder, Chirag Varun Shukla

Machine learning techniques paired with the availability of massive datasets dramatically enhance our ability to explore the chemical compound space by providing fast and accurate predictions of molecular properties. However, learning on large datasets is strongly limited by the availability of computational resources and can be infeasible in some scenarios. Moreover, the instances in the datasets may not yet be labelled and generating the labels can be costly, as in the case of quantum chemistry computations. Thus, there is a need to select small training subsets from large pools of unlabelled data points and to develop reliable ML methods that can effectively learn from small training sets. This work focuses on predicting the molecules atomization energy in the QM9 dataset. We investigate the advantages of employing domain knowledge-based data sampling methods for an efficient training set selection combined with informed ML techniques. In particular, we show how maximizing molecular diversity in the training set selection process increases the robustness of linear and nonlinear regression techniques such as kernel methods and graph neural networks. We also check the reliability of the predictions made by the graph neural network with a model-agnostic explainer based on the rate distortion explanation framework.

📄 PDF Abstract BibTeX arXiv:2306.10066

Code (0)

등록된 구현이 없습니다.

Tasks

Graph Neural Network

Methods 이 논문이 사용한 방법론

Graph Neural Network 설명 없음

Similar Papers 제목 키워드 기반

Diffusion Reconstruction-based Data Likelihood Estimation for Core-Set Selection

2025-11-24 · Mingyang Chen, Jiawei Du, Bo Huang, Yi Wang 외 arxiv

Existing core-set selection methods predominantly rely on heuristic scoring signals such as training dynamics or model uncertainty, lacking explicit modeling of data likelihood. This omission may hinder the constructed s…

Adaptive parameter-efficient fine-tuning via Hessian-informed subset selection

2025-05-18 · Shiyun Xu, Zhiqi Bu

Parameter-efficient fine-tuning (PEFT) is a highly effective approach for adapting large pre-trained models to downstream tasks with minimal computational overhead. At the core, PEFT methods freeze most parameters and on…

parameter-efficient fine-tuning

Diversity-Aware Adaptive Collocation for Physics-Informed Neural Networks via Sparse QUBO Optimization and Hybrid Coresets

2026-03-06 · Hadi Salloum, Maximilian Mifsud Bonici, Sinan Ibrahim, Pavel Osinenko 외 arxiv

Physics-Informed Neural Networks (PINNs) enforce governing equations by penalizing PDE residuals at interior collocation points, but standard collocation strategies - uniform sampling and residual-based adaptive refineme…

Optimality of Graphlet Screening in High Dimensional Variable Selection

2012-04-29 · Jiashun Jin, Cun-Hui Zhang, Qi Zhang

Consider a linear regression model where the design matrix X has n rows and p columns. We assume (a) p is much large than n, (b) the coefficient vector beta is sparse in the sense that only a small fraction of its coordi…

Variable SelectionVocal Bursts Intensity Prediction

Investigating the Impact of Data Selection Strategies on Language Model Performance

2025-01-07 · Jiayao Gu, Liting Chen, Yihong Li

Data selection is critical for enhancing the performance of language models, particularly when aligning training datasets with a desired target distribution. This study explores the effects of different data selection me…

Language ModelingLanguage Modelling