paper-with-me

Papers

TimeVista: Exploring and Exploiting Vision-Language Models as Judges for Time Series Forecasting

2026-06-15 · Zhi Chen, Yuxuan Wang, Jialong Wu, Yong Liu, Haoran Zhang, Xingjian Su, Jianmin Wang, Mingsheng Long arxiv

High-quality time series forecasting is pivotal for real-world decision-making. However, traditional point-wise metrics often fail to reveal complex temporal patterns and align poorly with human intuitive preferences. While the ''LLM-as-a-Judge'' paradigm has revolutionized text evaluation by providing flexible, human-aligned judgment, its application to time series remains largely unexplored. In this paper, we leverage Vision-Language Models (VLMs) as judges for time series forecasting, harnessing their ability to comprehend time series plots grounded in textual information. Specifically, we propose a novel framework integrating micro- and macro-level judgments informed by contextual information to evaluate time series forecasting. To this end, we introduce TimeVista, a comprehensive VLM-as-a-Judge benchmark comprising 5563 time series samples paired with detailed evaluation rubrics. Extensive meta-evaluations demonstrate that VLMs are highly reliable judges, achieving significantly higher consistency with human preferences than conventional metrics. Building upon our benchmark, we comprehensively assess recent Time Series Foundation Models (TSFMs) under the VLM-as-a-Judge paradigm. Our results demonstrate that VLMs serve as robust and interpretable judges, providing a comprehensive, human-aligned standard for evaluating time series models.

📄 PDF Abstract BibTeX arXiv:2606.16173

Code (0)

등록된 구현이 없습니다.

Tasks

Time Series Forecasting

Similar Papers 제목 키워드 기반

Don't Judge Code by Its Cover: Exploring Biases in LLM Judges for Code Evaluation

2025-05-22 · Jiwon Moon, Yerin Hwang, Dongryeol Lee, Taegwan Kang 외

With the growing use of large language models(LLMs) as evaluators, their application has expanded to code evaluation tasks, where they assess the correctness of generated code without relying on reference implementations…

Can Machines Imitate Humans? Integrative Turing Tests for Vision and Language Demonstrate a Narrowing Gap

2022-11-23 · Mengmi Zhang, Giorgia Dellaferrera, Ankur Sikarwar, Caishun Chen 외

As AI algorithms increasingly participate in daily activities, it becomes critical to ascertain whether the agents we interact with are human or not. To address this question, we turn to the Turing test and systematicall…

Image Captioningobject-detectionObject Detection

Exploiting LLM-as-a-Judge Disposition on Free Text Legal QA via Prompt Optimization

2026-04-22 · Mohamed Hesham Elganayni, Runsheng Chen, Sebastian Nagl, Matthias Grabmair arxiv

This work explores the role of prompt design and judge selection in LLM-as-a-Judge evaluations of free text legal question answering. We examine whether automatic task prompt optimization improves over human-centered des…

Question Answering

Malicious Repurposing of Open Science Artefacts by Using Large Language Models

2026-01-26 · Zahra Hashemi, Zhiqiang Zhong, Jun Pang, Wei Zhao arxiv

The rapid evolution of large language models (LLMs) has fuelled enthusiasm about their role in advancing scientific discovery, with studies exploring LLMs that autonomously generate and evaluate novel research ideas. How…

A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness

2026-02-04 · Leo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami 외 arxiv

Automated \enquote{LLM-as-a-Judge} frameworks have become the de facto standard for scalable evaluation across natural language processing. For instance, in safety evaluation, these judges are relied upon to evaluate har…

Adversarial Robustness