paper-with-me

홈 › Papers

PHM-Bench: A Domain-Specific Benchmarking Framework for Systematic Evaluation of Large Models in Prognostics and Health Management

2025-08-04 · Puyu Yang, Laifa Tao, Zijian Huang, Haifei Liu, Wenyan Cao, Hao Ji, Jianan Qiu, Qixuan Huang, Xuanyuan Su, Yuhang Xie, Jun Zhang, Shangyu Li, Chen Lu, Zhixuan Lian arxiv

With the rapid advancement of generative artificial intelligence, large language models (LLMs) are increasingly adopted in industrial domains, offering new opportunities for Prognostics and Health Management (PHM). These models help address challenges such as high development costs, long deployment cycles, and limited generalizability. However, despite the growing synergy between PHM and LLMs, existing evaluation methodologies often fall short in structural completeness, dimensional comprehensiveness, and evaluation granularity. This hampers the in-depth integration of LLMs into the PHM domain. To address these limitations, this study proposes PHM-Bench, a novel three-dimensional evaluation framework for PHM-oriented large models. Grounded in the triadic structure of fundamental capability, core task, and entire lifecycle, PHM-Bench is tailored to the unique demands of PHM system engineering. It defines multi-level evaluation metrics spanning knowledge comprehension, algorithmic generation, and task optimization. These metrics align with typical PHM tasks, including condition monitoring, fault diagnosis, RUL prediction, and maintenance decision-making. Utilizing both curated case sets and publicly available industrial datasets, our study enables multi-dimensional evaluation of general-purpose and domain-specific models across diverse PHM tasks. PHM-Bench establishes a methodological foundation for large-scale assessment of LLMs in PHM and offers a critical benchmark to guide the transition from general-purpose to PHM-specialized models.

📄 PDF Abstract BibTeX arXiv:2508.02490

Code (0)

등록된 구현이 없습니다.

Tasks

Fault Diagnosis

Similar Papers 제목 키워드 기반

Enterprise Benchmarks for Large Language Model Evaluation

2024-10-11 · Bing Zhang, Mikio Takeuchi, Ryo Kawahara, Shubhi Asthana 외

The advancement of large language models (LLMs) has led to a greater challenge of having a rigorous and systematic evaluation of complex tasks performed, especially in enterprise applications. Therefore, LLMs need to be …

BenchmarkingLanguage Model EvaluationLanguage ModelingLanguage Modelling+2

Compressed Video Quality Enhancement: Classifying and Benchmarking over Standards

2025-09-12 · Xiem HoangVan, Dang BuiDinh, Sang NguyenQuang, Wen-Hsiao Peng arxiv

Compressed video quality enhancement (CVQE) is crucial for improving user experience with lossy video codecs like H.264/AVC, H.265/HEVC, and H.266/VVC. While deep learning based CVQE has driven significant progress, exis…

GenBench: A Benchmarking Suite for Systematic Evaluation of Genomic Foundation Models

2024-06-01 · Zicheng Liu, Jiahui Li, Siyuan Li, Zelin Zang 외

The Genomic Foundation Model (GFM) paradigm is expected to facilitate the extraction of generalizable representations from massive genomic data, thereby enabling their application across a spectrum of downstream applicat…

Benchmarking

Towards Large Scale Automated Algorithm Design by Integrating Modular Benchmarking Frameworks

2021-02-12 · Amine Aziz-Alaoui, Carola Doerr, Johann Dreo

We present a first proof-of-concept use-case that demonstrates the efficiency of interfacing the algorithm framework ParadisEO with the automated algorithm configuration tool irace and the experimental platform IOHprofil…

Benchmarking

Benchmarking Adversarial Robustness to Bias Elicitation in Large Language Models: Scalable Automated Assessment with LLM-as-a-Judge

2025-04-10 · Riccardo Cantini, Alessio Orsino, Massimo Ruggiero, Domenico Talia

Large Language Models (LLMs) have revolutionized artificial intelligence, driving advancements in machine translation, summarization, and conversational agents. However, their increasing integration into critical societa…

Adversarial RobustnessBenchmarkingFairnessMachine Translation