paper-with-me

Papers

TS-MLLM: A Multi-Modal Large Language Model-based Framework for Industrial Time-Series Big Data Analysis

2026-03-08 · Haiteng Wang, Yikang Li, Yunfei Zhu, Jingheng Yan, Lei Ren, Laurence T. Yang arxiv

Accurate analysis of industrial time-series big data is critical for the Prognostics and Health Management (PHM) of industrial equipment. While recent advancements in Large Language Models (LLMs) have shown promise in time-series analysis, existing methods typically focus on single-modality adaptations, failing to exploit the complementary nature of temporal signals, frequency-domain visual representations, and textual knowledge information. In this paper, we propose TS-MLLM, a unified multi-modal large language model framework designed to jointly model temporal signals, frequency-domain images, and textual domain knowledge. Specifically, we first develop an Industrial time-series Patch Modeling branch to capture long-range temporal dynamics. To integrate cross-modal priors, we introduce a Spectrum-aware Vision-Language Model Adaptation (SVLMA) mechanism that enables the model to internalize frequency-domain patterns and semantic context. Furthermore, a Temporal-centric Multi-modal Attention Fusion (TMAF) mechanism is designed to actively retrieve relevant visual and textual cues using temporal features as queries, ensuring deep cross-modal alignment. Extensive experiments on multiple industrial benchmarks demonstrate that TS-MLLM significantly outperforms state-of-the-art methods, particularly in few-shot and complex scenarios. The results validate our framework's superior robustness, efficiency, and generalization capabilities for industrial time-series prediction.

📄 PDF Abstract BibTeX arXiv:2603.07572

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Orchestrate Multimodal Data with Batch Post-Balancing to Accelerate Multimodal Large Language Model Training

2025-03-31 · Yijie Zheng, Bangjun Xiao, Lei Shi, Xiaoyang Li 외

Multimodal large language models (MLLMs), such as GPT-4o, are garnering significant attention. During the exploration of MLLM training, we identified Modality Composition Incoherence, a phenomenon that the proportion of …

GPULanguage ModelingLanguage ModellingLarge Language Model+1

LLaVA-KD: A Framework of Distilling Multimodal Large Language Models

2024-10-21 · Yuxuan Cai, Jiangning Zhang, Haoyang He, Xinwei He 외

The success of Large Language Models (LLM) has led researchers to explore Multimodal Large Language Models (MLLM) for unified visual and linguistic understanding. However, the increasing model size and computational comp…

DreamLLM: Synergistic Multimodal Comprehension and Creation

2023-09-20 · Runpei Dong, Chunrui Han, Yuang Peng, Zekun Qi 외

This paper presents DreamLLM, a learning framework that first achieves versatile Multimodal Large Language Models (MLLMs) empowered with frequently overlooked synergy between multimodal comprehension and creation. DreamL…

multimodal generationVisual Question AnsweringZero-Shot LearningZero-Shot Text-to-Image Generation

PlaM: Training-Free Plateau-Guided Model Merging for Better Visual Grounding in MLLMs

2026-01-12 · Zijing Wang, Yongkang Liu, Mingyang Wang, Ercong Nie 외 arxiv

Multimodal Large Language Models (MLLMs) rely on strong linguistic reasoning inherited from their base language models. However, multimodal instruction fine-tuning paradoxically degrades this text's reasoning capability,…

Visual Grounding

E5-V: Universal Embeddings with Multimodal Large Language Models

2024-07-17 · Ting Jiang, Minghui Song, Zihan Zhang, Haizhen Huang 외

Multimodal large language models (MLLMs) have shown promising advancements in general visual and language understanding. However, the representation of multimodal information using MLLMs remains largely unexplored. In th…