paper-with-me

홈 › Papers

VTBench: A Multimodal Framework for Time-Series Classification with Chart-Based Representations

2026-04-29 · Madhumitha Venkatesan, Xuyang Chen, Dongyu Liu arxiv

Time-series classification (TSC) has advanced significantly with deep learning, yet most models rely solely on raw numerical inputs, overlooking alternative representations. While texture-based encodings such as Gramian Angular Fields (GAF) and Recurrence Plots (RP) convert time series into 2D images, they often require heavy preprocessing and yield less intuitive representations. In contrast, chart-based visualizations offer more interpretable alternatives and show promise in specific domains; however, their effectiveness remains underexplored, with limited systematic evaluation across chart types, visual encoding choices, and datasets. In this work, we introduce VTBench, a systematic and extensible framework that re-examines TSC through multimodal fusion of raw sequences and chart-based visualizations. VTBench generates lightweight, human-interpretable plots -- line, area, bar, and scatter, providing complementary views of the same signal. We develop a modular architecture supporting multiple fusion strategies, including single-chart visual-numerical fusion, multi-chart visual fusion, and full multimodal fusion with raw inputs. Through experiments across 31 UCR datasets, we show that: (1) chart-only models are competitive in selected settings, particularly on smaller datasets; (2) combining multiple chart types can improve accuracy by capturing complementary visual cues; and (3) multimodal models improve or maintain performance when visual features provide non-redundant information, but may degrade accuracy when they introduce redundancy. We further distill practical guidelines for selecting chart types, fusion strategies, and configurations. VTBench establishes a unified foundation for interpretable and effective multimodal time-series classification.

📄 PDF Abstract BibTeX arXiv:2604.27259

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

VTBench: Comprehensive Benchmark Suite Towards Real-World Virtual Try-on Models

2025-05-26 · Hu Xiaobin, Liang Yujie, Luo Donghao, Peng Xu 외

While virtual try-on has achieved significant progress, evaluating these models towards real-world scenarios remains a challenge. A comprehensive benchmark is essential for three key reasons:(1) Current metrics inadequat…

Occlusion HandlingVirtual Try-on

SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation

2026-08-19 · Bowen Jing, Mingxin Wang, Ruiyang Hao, Chenchen Ge 외 arxiv

Physical interaction quality is central to deformable-object manipulation, yet most benchmarks evaluate task success alone. A policy may complete the task while allowing slip or causing excessive compression. A primary b…

SoftVTBench: A Safety-Aware Visuo-Tactile Benchmark for Physically Constrained Robotic Manipulation of Deformable Objects (Early Version)

2026-07-05 · Bowen Jing, Mingxin Wang, Ruiyang Hao, Chenchen Ge 외 arxiv

Deformable object manipulation poses challenges beyond task completion: successful execution must also maintain safe physical interaction, holding the object stably without slip or drop while avoiding excessive deformati…

InstructTime++: Time Series Classification with Multimodal Language Modeling via Implicit Feature Enhancement

2026-01-21 · Mingyue Cheng, Xiaoyu Tao, Huajian Zhang, Qi Liu 외 arxiv

Most existing time series classification methods adopt a discriminative paradigm that maps input sequences directly to one-hot encoded class labels. While effective, this paradigm struggles to incorporate contextual feat…

Time Series ClassificationImage Captioning

Pixel-Wise Multimodal Contrastive Learning for Remote Sensing Images

2026-01-07 · Leandro Stival, Ricardo da Silva Torres, Helio Pedrini arxiv

Satellites continuously generate massive volumes of data, particularly for Earth observation, including satellite image time series (SITS). However, most deep learning models are designed to process either entire images …

Contrastive Learning