paper-with-me

홈 › Papers

Lessons from the trenches on evaluating machine-learning systems in materials science

2025-03-13 · Nawaf Alampara, Mara Schilling-Wilhelmi, Kevin Maik Jablonka

Measurements are fundamental to knowledge creation in science, enabling consistent sharing of findings and serving as the foundation for scientific discovery. As machine learning systems increasingly transform scientific fields, the question of how to effectively evaluate these systems becomes crucial for ensuring reliable progress. In this review, we examine the current state and future directions of evaluation frameworks for machine learning in science. We organize the review around a broadly applicable framework for evaluating machine learning systems through the lens of statistical measurement theory, using materials science as our primary context for examples and case studies. We identify key challenges common across machine learning evaluation such as construct validity, data quality issues, metric design limitations, and benchmark maintenance problems that can lead to phantom progress when evaluation frameworks fail to capture real-world performance needs. By examining both traditional benchmarks and emerging evaluation approaches, we demonstrate how evaluation choices fundamentally shape not only our measurements but also research priorities and scientific progress. These findings reveal the critical need for transparency in evaluation design and reporting, leading us to propose evaluation cards as a structured approach to documenting measurement choices and limitations. Our work highlights the importance of developing a more diverse toolbox of evaluation techniques for machine learning in materials science, while offering insights that can inform evaluation practices in other scientific domains where similar challenges exist.

📄 PDF Abstract BibTeX arXiv:2503.10837

Code (0)

등록된 구현이 없습니다.

Tasks

scientific discovery

Similar Papers 제목 키워드 기반

Lessons from the Trenches on Reproducible Evaluation of Language Models

2024-05-23 · Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao 외

Effective evaluation of language models remains an open challenge in NLP. Researchers and engineers face methodological issues such as the sensitivity of models to evaluation setup, difficulty of proper comparisons acros…

Language Model EvaluationLanguage ModelingLanguage Modelling

User-centered & Robust NLP OSS: Lessons Learned from Developing & Maintaining RSMTool

2020-11-01 · EMNLP (NLPOSS) 2020 11 · Nitin Madnani, Anastassia Loukina

For the last 5 years, we have developed and maintained RSMTool – an open-source tool for evaluating NLP systems that automatically score written and spoken responses. RSMTool is designed to be cross-disciplinary, borrowi…

Low-Precision Hardware Architectures Meet Recommendation Model Inference at Scale

2021-05-26 · Zhaoxia, Deng, Jongsoo Park, Ping Tak Peter Tang 외

Tremendous success of machine learning (ML) and the unabated growth in ML model complexity motivated many ML-specific designs in both CPU and accelerator architectures to speed up the model inference. While these archite…

CPURecommendation Systems

T-RAG: Lessons from the LLM Trenches

2024-02-12 · Masoomali Fatehkia, Ji Kim Lucas, Sanjay Chawla

Large Language Models (LLM) have shown remarkable language capabilities fueling attempts to integrate them into applications across a wide range of domains. An important application area is question answering over privat…

Question AnsweringRAGRetrieval-augmented Generation

First Lessons Learned of an Artificial Intelligence Robotic System for Autonomous Coarse Waste Recycling Using Multispectral Imaging-Based Methods

2025-01-23 · Timo Lange, Ajish Babu, Philipp Meyer, Matthis Keppner 외

Current disposal facilities for coarse-grained waste perform manual sorting of materials with heavy machinery. Large quantities of recyclable materials are lost to coarse waste, so more effective sorting processes must b…

Material Classificationobject-detectionObject Detection