paper-with-me

홈 › Papers

STEP: Structured Training and Evaluation Platform for benchmarking trajectory prediction models

2025-09-18 · Julian F. Schumann, Anna Mészáros, Jens Kober, Arkady Zgonnikov arxiv

While trajectory prediction plays a critical role in enabling safe and effective path-planning in automated vehicles, standardized practices for evaluating such models remain underdeveloped. Recent efforts have aimed to unify dataset formats and model interfaces for easier comparisons, yet existing frameworks often fall short in supporting heterogeneous traffic scenarios, joint prediction models, or user documentation. In this work, we introduce STEP -- a new benchmarking framework that addresses these limitations by providing a unified interface for multiple datasets, enforcing consistent training and evaluation conditions, and supporting a wide range of prediction models. We demonstrate the capabilities of STEP in a number of experiments which reveal 1) the limitations of widely-used testing procedures, 2) the importance of joint modeling of agents for better predictions of interactions, and 3) the vulnerability of current state-of-the-art models against both distribution shifts and targeted attacks by adversarial agents. With STEP, we aim to shift the focus from the ``leaderboard'' approach to deeper insights about model behavior and generalization in complex multi-agent settings.

📄 PDF Abstract BibTeX arXiv:2509.14801

Code (0)

등록된 구현이 없습니다.

Tasks

Trajectory Prediction

Similar Papers 제목 키워드 기반

The Design and Implementation of a Scalable DL Benchmarking Platform

2019-11-19 · Cheng Li, Abdul Dakkak, JinJun Xiong, Wen-mei Hwu

The current Deep Learning (DL) landscape is fast-paced and is rife with non-uniform models, hardware/software (HW/SW) stacks, but lacks a DL benchmarking platform to facilitate evaluation and comparison of DL innovations…

Benchmarking

Arena-Rosnav 2.0: A Development and Benchmarking Platform for Robot Navigation in Highly Dynamic Environments

2023-02-20 · Linh Kästner, Reyk Carstens, Huajian Zeng, Jacek Kmiecik 외

Following up on our previous works, in this paper, we present Arena-Rosnav 2.0 an extension to our previous works Arena-Bench and Arena-Rosnav, which adds a variety of additional modules for developing and benchmarking r…

BenchmarkingRobot Navigation

BenGER Platform: A Collaborative Web Platform for End-to-End Benchmarking of German Legal Tasks

2026-04-15 · Sebastian Nagl, Matthias Grabmair arxiv

Evaluating large language models (LLMs) for legal reasoning requires workflows that span task design, expert annotation, model execution, and metric-based evaluation. In practice, these steps are split across platforms a…

Legal Reasoning

REPLAB: A Reproducible Low-Cost Arm Benchmark Platform for Robotic Learning

2019-05-17 · Brian Yang, Jesse Zhang, Vitchyr Pong, Sergey Levine 외

Standardized evaluation measures have aided in the progress of machine learning approaches in disciplines such as computer vision and machine translation. In this paper, we make the case that robotic learning would also …

BenchmarkingDeep Reinforcement LearningMachine TranslationReinforcement Learning+1

BMOBench: Black-Box Multi-Objective Optimization Benchmarking Platform

2016-05-23 · Abdullah Al-Dujaili, S. Suresh

This document briefly describes the Black-Box Multi-Objective Optimization Benchmarking (BMOBench) platform. It presents the test problems, evaluation procedure, and experimental setup. To this end, the BMOBench is demon…

Benchmarking