paper-with-me

Papers

SSP: An Event-Matched Syn2Sim2Phy Cross-Domain Evaluation Framework for Autonomous Driving VLA Models

2026-08-14 · Haojie Feng, Peizhi Zhang, Xinrui Zhang, Zhuoren Li, Junpeng Huang, Xiurong Wang, Dongxiao Yin, Yuxiang Zhang, Junfan Zhu, Lu Xiong arxiv

Vision-language-action (VLA) models for autonomous driving jointly produce scene interpretation, language-based reasoning, and driving trajectories. Existing evaluations often use independently selected synthetic, simulated, and physical data, so measured performance gaps can be confounded by changes in scenario content rather than genuine domain sensitivity. We propose SSP (Synthetic-Simulation-Physical), an event-matched Syn2Sim2Phy evaluation framework that anchors cross-domain comparison to the same safety-critical interaction. Starting from a synthetic long-tail video, SSP builds a validated event specification that preserves road topology, participant roles, relative motion, conflict evolution, passing order, response constraints, and event phases. Platform-specific realizations are then constructed in CARLA and on a closed proving ground and are evaluated only after transfer audits confirm preservation of mandatory event properties. SSP maps heterogeneous outputs from OpenEMMA, LLaViDA, and Alpamayo-R1 into common semantic slots and a 1 s trajectory window to assess output validity, semantic accuracy, critical-interaction recognition, trajectory quality, and risk response. Across Cut-in and vulnerable-road-user crossing cases, the macro-averaged Integrated VLA Capability Scores are 0.259, 0.291, and 0.325 in the Synthetic, Simulation, and Physical domains, respectively, while the best domain varies by scenario. Alpamayo-R1, OpenEMMA, and LLaViDA obtain scores of 0.405, 0.338, and 0.131. SSP provides a reproducible scene-transfer chain and an evidence-qualified evaluation of VLA behavior without assuming that the Physical domain is universally superior.

📄 PDF Abstract BibTeX arXiv:2608.14024

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

Motion-Based Tokenization for Cross-Dataset Egocentric Gaze Modeling

2026-08-24 · Virmarie Maquiling, Zhuojiang Cai, Enkelejda Kasneci arxiv

Gaze is increasingly used as an input signal for vision and multimodal models, yet no consensus exists on how to represent it across datasets. Raw traces preserve detail but are noisy and device-dependent, while coarse e…

Does Your Wildfire Prediction Model Actually Work, or Just Score Well?

2026-05-14 · Yangshuang Xu, Yuyang Dai, Liling Chang, Qi Wang 외 arxiv

Wildfire prediction is important for early warning and resource allocation, yet existing Earth foundation models (Earth FMs) are pretrained for general atmospheric and geophysical objectives rather than wildfire forecast…

Fishr: Invariant Gradient Variances for Out-of-Distribution Generalization

2021-09-07 · Alexandre Rame, Corentin Dancette, Matthieu Cord

Learning robust models that generalize well under changes in the data distribution is critical for real-world applications. To this end, there has been a growing surge of interest to learn simultaneously from multiple tr…

Domain GeneralizationOut-of-Distribution Generalization

On the Signal Processing Operations in LIGO signals

2018-03-12

This article analyzes the data for the five gravitational wave (GW) events detected in Hanford(H1), Livingston(L1) and Virgo(V1) detectors by the LIGO collaboration. It is shown that GW170814, GW170817, GW151226 and GW17…

Domain Mismatch Doesn’t Always Prevent Cross-lingual Transfer Learning

2022-06-01 · LREC 2022 6 · Daniel Edmiston, Phillip Keung, Noah A. Smith

Cross-lingual transfer learning without labeled target language data or parallel text has been surprisingly effective in zero-shot cross-lingual classification, question answering, unsupervised machine translation, etc. …

Bilingual Lexicon InductionCross-Lingual TransferMachine TranslationQuestion Answering+4