paper-with-me

홈 › Papers

Full-data accuracy with fewer labels for training and fine-tuning machine-learning force fields

2026-07-16 · Sheng Bi, Yi-Ze Wang, Jun Cheng arxiv

Machine-learning force fields (MLFFs) are reliable only near their training distribution, making efficient construction of diverse training sets a major bottleneck for both train-from-scratch and foundation fine-tuning workflows. Active learning can reduce this cost, but standard model-committee uncertainty is impractical for foundation MLFFs because each committee member requires a separate fine-tuning run. We present an active-learning workflow based on last-layer-projection regression (LLPR), a forward-pass-cheap per-configuration uncertainty estimator. Across molecular, condensed-phase, and electrolyte systems, LLPR identifies compact, high-value training sets that recover full-data accuracy using only a small fraction of electronic-structure labels. In foundation-model fine-tuning, LLPR-selected configurations reach the full-pool fine-tuning ceiling with substantially fewer labels than random selection. In iterative electrolyte fine-tuning, LLPR detects unphysical local coordination before DFT labelling, provides an absolute force-error threshold, and enables automatic termination of the learning loop. The resulting models reproduce reference density and ion-coordination structure, providing a scalable uncertainty-quantification strategy across MLFF training regimes.

📄 PDF Abstract BibTeX arXiv:2607.14486

Code (1)

exopoiesis/arxiv-radar-physics ★ 1

Tasks

Active Learning

Similar Papers 제목 키워드 기반

What If We Only Use Real Datasets for Scene Text Recognition? Toward Scene Text Recognition With Fewer Labels

2021-03-07 · CVPR 2021 1 · Jeonghun Baek, Yusuke Matsui, Kiyoharu Aizawa

Scene text recognition (STR) task has a common practice: All state-of-the-art STR models are trained on large synthetic data. In contrast to this practice, training STR models only on fewer real labels (STR with fewer la…

Data AugmentationScene Text Recognition

Weakly Supervised Semantic Point Cloud Segmentation: Towards 10x Fewer Labels

2020-06-01 · CVPR 2020 6 · Xun Xu, Gim Hee Lee

Point cloud analysis has received much attention recently; and segmentation is one of the most important tasks. The success of existing approaches is attributed to deep network design and large amount of labelled trainin…

Point Cloud SegmentationSegmentationWeakly Supervised 3D Point Cloud Segmentation

Weakly Supervised Semantic Point Cloud Segmentation:Towards 10X Fewer Labels

2020-04-08 · Xun Xu, Gim Hee Lee

Point cloud analysis has received much attention recently; and segmentation is one of the most important tasks. The success of existing approaches is attributed to deep network design and large amount of labelled trainin…

Point Cloud SegmentationSegmentation

Are Fewer Labels Possible for Few-shot Learning?

2020-12-10 · Suichan Li, Dongdong Chen, Yinpeng Chen, Lu Yuan 외

Few-shot learning is challenging due to its very limited data and labels. Recent studies in big transfer (BiT) show that few-shot learning can greatly benefit from pretraining on large scale labeled dataset in a differen…

ClusteringFew-Shot Learning

Pseudo-Label Generation and Various Data Augmentation for Semi-Supervised Hyperspectral Object Detection

2022-10-01 · Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops 2022 10 · Jun Yu, Liwen Zhang, Shenshen Du, Hao Chang 외

Semi-supervised learning is a highly researched problem, but existing semi-supervised object detection frameworks are based on RGB images, and existing pre-trained models cannot be used for hyperspectral images. To overc…

Data Augmentationobject-detectionObject DetectionPseudo Label+1