paper-with-me

홈 › Papers

HUGO-CS: A Hybrid-Labeled, Uncertainty-Aware, General-Purpose, Observational Dataset for Cold Spray

2026-05-05 · Stephen Price, Kyle Miller, Marco Musto, Kenneth Kroenlein, James Saal, Kyle Tsaknopoulos, Elke A. Rundensteiner, Danielle L. Cote arxiv

Cold spraying is an increasingly common approach for repairing and manufacturing components due to its solid-state manufacturing capabilities. However, process optimization remains difficult due to many interdependent parameters and the lack of large-scale, machine-readable data to support modeling. While the scientific literature contains many relevant experiments, results are inconsistently reported (often in tables and figures) and use non-uniform units, limiting utilization at scale. To address these limitations, this work presents HUGO-CS, a literature-derived dataset of 4,383 cold-spray experiments with 144 features from 1,124 sources, exceeding the previous largest dataset (137 samples) by 30x. With completely manual extraction requiring an average of 91 minutes per document, this work designs and leverages a Hybrid-labeled, Uncertainty-aware, General-purpose, Observational extraction framework, called HUGO, to support this extraction. HUGO combines automated LLM-based labeling with targeted manual label refinement to handle this experimental result extraction process from scientific literature. To balance labeling efficiency with extraction accuracy, HUGO introduces a Hierarchical Risk Mitigation (HRM) to route LLM outputs with a high risk of potential errors for manual review, while retaining low-risk records as auto-labeled. Lastly, HUGO post-processing consolidates categorical descriptors, maps reported feedstock chemistries into structured continuous compositions, and normalizes units across sources. Of the 4,383 reported experiments, 1,765 are hand-labeled, providing a high-quality labeled subset for benchmarking, error analysis, and higher-fidelity data points. All code to replicate this work, along with the complete HUGO-CS dataset, are released under a CC-BY license at https://github.com/sprice134/HUGO.

📄 PDF Abstract BibTeX arXiv:2605.04257

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Physics-constrained Gaussian Processes for Predicting Shockwave Hugoniot Curves

2026-01-10 · George D. Pasparakis, Himanshu Sharma, Rushik Desai, Chunyu Li 외 arxiv

A physics-constrained Gaussian Process regression framework is developed for predicting shocked material states and their associated uncertainties along the Hugoniot curve using data from a small number of shockwave simu…

Gaussian Processes

Learning thermodynamically constrained equations of state with uncertainty

2023-06-29 · Himanshu Sharma, Jim A. Gaffney, Dimitrios Tsapetis, Michael D. Shields

Numerical simulations of high energy-density experiments require equation of state (EOS) models that relate a material's thermodynamic state variables -- specifically pressure, volume/density, energy, and temperature. EO…

GPRUncertainty Quantification

VIOLA: Towards Video In-Context Learning with Minimal Annotations

2026-01-22 · Ryo Fujii, Hideo Saito, Ryo Hachiuma arxiv

Generalizing Multimodal Large Language Models (MLLMs) to novel video domains is essential for real-world deployment but remains challenging due to the scarcity of labeled data. While In-Context Learning (ICL) offers a tr…

Density Estimation

Uncertainty-aware transfer across tasks using hybrid model-based successor feature reinforcement learning

2023-10-16 · Parvin Malekzadeh, Ming Hou, Konstantinos N. Plataniotis

Sample efficiency is central to developing practical reinforcement learning (RL) for complex and large-scale decision-making problems. The ability to transfer and generalize knowledge gained from previous experiences to …

Decision MakingReinforcement Learning (RL)Transfer Learning

To Label or Not to Label: Hybrid Active Learning for Neural Machine Translation

2024-03-14 · Abdul Hameed Azeemi, Ihsan Ayyub Qazi, Agha Ali Raza

Active learning (AL) techniques reduce labeling costs for training neural machine translation (NMT) models by selecting smaller representative subsets from unlabeled data for annotation. Diversity sampling techniques sel…

Active LearningDiversityDomain AdaptationMachine Translation+3