VSLLaVA: a pipeline of large multimodal foundation model for industrial vibration signal analysis
Large multimodal foundation models have been extensively utilized for image recognition tasks guided by instructions, yet there remains a scarcity of domain expertise in industrial vibration signal analysis. This paper presents a pipeline named VSLLaVA that leverages a large language model to integrate expert knowledge for identification of signal parameters and diagnosis of faults. Within this pipeline, we first introduce an expert rule-assisted signal generator. The generator merges signal provided by vibration analysis experts with domain-specific parameter identification and fault diagnosis question-answer pairs to build signal-question-answer triplets. Then we use these triplets to apply low-rank adaptation methods for fine-tuning the linear layers of the Contrastive Language-Image Pretraining (CLIP) and large language model, injecting multimodal signal processing knowledge. Finally, the fine-tuned model is assessed through the combined efforts of large language model and expert rules to evaluate answer accuracy and relevance, which showcases enhanced performance in identifying, analyzing various signal parameters, and diagnosing faults. These enhancements indicate the potential of this pipeline to build a foundational model for future industrial signal analysis and monitoring.
Code (0)
등록된 구현이 없습니다.
Tasks
Fault DiagnosisLanguage ModelingLanguage ModellingLarge Language ModelSimilar Papers 제목 키워드 기반
Industrial Language-Image Dataset (ILID): Adapting Vision Foundation Models for Industrial Settings
In recent years, the upstream of Large Language Models (LLM) has also encouraged the computer vision community to work on substantial multimodal datasets and train models on a scale in a self-/semi-supervised manner, res…
Transfer LearningTowards Open-Vocabulary Industrial Defect Understanding with a Large-Scale Multimodal Dataset
We present IMDD-1M, the first large-scale Industrial Multimodal Defect Dataset comprising 1,000,000 aligned image-text pairs, designed to advance multimodal learning for manufacturing and quality inspection. IMDD-1M cont…
Xuanwu: Evolving General Multimodal Models into an Industrial-Grade Foundation for Content Ecosystems
In recent years, multimodal large models have continued to improve on general benchmarks. However, in real-world content moderation and adversarial settings, mainstream models still suffer from degraded generalization an…
Towards General Industrial Intelligence: A Survey of Continual Large Models in Industrial IoT
Industrial AI is transitioning from traditional deep learning models to large-scale transformer-based architectures, with the Industrial Internet of Things (IIoT) playing a pivotal role. IIoT evolves from a simple data p…
Cloud ComputingContinual LearningEdge-computingAutomatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection
Identifying defects and anomalies in industrial products is a critical quality control task. Traditional manual inspection methods are slow, subjective, and error-prone. In this work, we propose a novel zero-shot trainin…
Anomaly DetectionImage-text matchingLanguage ModelingLanguage Modelling+4