paper-with-me

Papers

SPHINX: First Explain, Then Explore

2026-06-16 · Nguyen Do, Tue M. Cao, Tien Van Do, András Hajdu, Tamás Bérczes, My T. Thai arxiv

Generating adversarial driving scenarios is critical for evaluating and improving autonomous vehicle decision-making systems in simulation. Recent approaches rely primarily on the prior knowledge of Large Language Models and Vision-Language Models to generate driving scenarios procedurally. We argue that adversarial scenes should be generated based on the failure diagnosis (e.g., indecisiveness, multi-frame inconsistency) of the driving policy to specifically address the policy's weaknesses instead of relying on prior assumptions. In this paper, we propose SPHINX, a closed-loop framework for adversarial scenario synthesis guided by a simple principle: first explain, then explore. Beyond blindly exploring the scenario space, SPHINX leverages explainable artificial intelligence methods to analyze the policy, identifying key visual concepts and their influence on policy outputs, and the uncertainty of the decisions. Given the interpretable evidence extracted from the policy's own decision process, we use a vision language model to rationalize and criticize failure modes of the current policy. These critics are then used to generate targeted adversarial scenarios for policy retraining and improvement. We demonstrate that SPHINX can highlight an interpretable account of policy failures while other adversarial scene generation cannot. Across the evaluated benchmarks and test suites, SPHINX can be applied to diverse state-of-the-art autonomous vehicle architectures and yields consistent robustness improvements over existing scenario-generation methods.

📄 PDF Abstract BibTeX arXiv:2606.17482

Code (0)

등록된 구현이 없습니다.

Tasks

Scene Generation

Similar Papers 제목 키워드 기반

sPhinX: Sample Efficient Multilingual Instruction Fine-Tuning Through N-shot Guided Prompting

2024-07-13 · Sanchit Ahuja, Kumar Tanmay, Hardik Hansrajbhai Chauhan, Barun Patra 외

Despite the remarkable success of LLMs in English, there is a significant gap in performance in non-English languages. In order to address this, we introduce a novel recipe for creating a multilingual synthetic instructi…

Machine TranslationQuestion AnsweringReading Comprehension

VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL

2025-05-29 · Yichen Feng, Zhangchen Xu, Fengqing Jiang, Yuetai Li 외

Vision language models (VLMs) are expected to perform effective multimodal reasoning and make logically coherent decisions, which is critical to tasks such as diagram understanding and spatial problem solving. However, c…

Arithmetic ReasoningImage GenerationLogical ReasoningMultimodal Reasoning

基於Sphinx 可快速個人化行動數字語音辨識系統 (Quickly Personalizable Mobile Digit Speech Recognition System Based on Sphinx) [In Chinese]

2013-10-01 · ROCLINGIJCLCLP 2013 10 · Tsung-Peng Yen, Chia-Ping Chen
speech-recognitionSpeech Recognition

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models

2023-11-13 · Ziyi Lin, Chris Liu, Renrui Zhang, Peng Gao 외

We present SPHINX, a versatile multi-modal large language model (MLLM) with a joint mixing of model weights, tuning tasks, and visual embeddings. First, for stronger vision-language alignment, we unfreeze the large langu…

Described Object DetectionLanguage ModelingLanguage ModellingLarge Language Model+4

Sphinx: Efficiently Serving Novel View Synthesis using Regression-Guided Selective Refinement

2025-11-24 · Yuchen Xia, Souvik Kundu, Mosharaf Chowdhury, Nishil Talati arxiv

Novel View Synthesis (NVS) is the task of generating new images of a scene from viewpoints that were not part of the original input. Diffusion-based NVS can generate high-quality, temporally consistent images, however, r…

Novel View Synthesis