paper-with-me

홈 › Papers

ScenePilot-4K: A Large-Scale First-Person Dataset and Benchmark for Vision-Language Models in Autonomous Driving

2026-01-27 · Yujin Wang, Yutong Zheng, Wenxian Fan, Tianyi Wang, Hongqing Chu, Li Zhang, Bingzhao Gao, Daxin Tian, Jianqiang Wang, Hong Chen arxiv

In this paper, we introduce ScenePilot-4K, a large-scale first-person dataset for safety-aware vision-language learning and evaluation in autonomous driving. Built from public online driving videos, ScenePilot-4K contains 3,847 hours of video and 27.7M front-view frames spanning 63 countries/regions and 1,210 cities. It jointly provides scene-level natural-language descriptions, risk assessment labels, key-participant annotations, ego trajectories, and camera parameters through a unified multi-stage annotation pipeline. Building on this dataset, we establish ScenePilot-Bench, a standardized benchmark that evaluates vision-language models along four complementary axes: scene understanding, spatial perception, motion planning, and GPT-based semantic alignment. The benchmark includes fine-grained metrics and geographic generalization settings that expose model robustness under cross-region and cross-traffic domain shifts. Baseline results on representative open-source and proprietary vision-language models show that current models remain competitive in high-level scene semantics but still exhibit substantial limitations in geometry-aware perception and planning-oriented reasoning. Beyond the released dataset itself, the proposed annotation pipeline serves as a reusable and extensible recipe for scalable dataset construction from public Internet driving videos. The codes and supplementary materials are available at: https://github.com/yjwangtj/ScenePilot-4K, with the dataset available at https://huggingface.co/datasets/larswangtj/ScenePilot-4K.

📄 PDF Abstract BibTeX arXiv:2601.19582

Code (0)

등록된 구현이 없습니다.

Tasks

Scene UnderstandingAutonomous DrivingMotion Planning

Similar Papers 제목 키워드 기반

ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation

2026-08-31 · Jiawei Zhang, Hongsong Wang, Pan Zhou arxiv

Text-driven 3D indoor scene generation has advanced from dataset-bound layout modeling to open-vocabulary synthesis with large language and vision-language models. Yet existing methods remain limited: one-pass generators…

Scene Generation

ScenePilot: Controllable Boundary-Driven Critical Scenario Generation for Autonomous Driving

2026-05-20 · Qiyu Ruan, Yuxuan Wang, He Li, Zhenning Li 외 arxiv

Safety-critical scenarios are central to evaluating autonomous driving systems, yet their rarity in naturalistic logs makes simulation-based stress testing indispensable. Most scenario generation methods treat surroundin…

Reinforcement LearningAutonomous Driving

FT-HID: A Large Scale RGB-D Dataset for First and Third Person Human Interaction Analysis

2022-09-21 · Zihui Guo, Yonghong Hou, Pichao Wang, Zhimin Gao 외

Analysis of human interaction is one important research topic of human motion analysis. It has been studied either using first person vision (FPV) or third person vision (TPV). However, the joint learning of both types o…

Action AnalysisAction Recognition

Unsupervised Pre-training for Person Re-identification

2020-12-07 · CVPR 2021 1 · Dengpan Fu, Dongdong Chen, Jianmin Bao, Hao Yang 외

In this paper, we present a large scale unlabeled person re-identification (Re-ID) dataset "LUPerson" and make the first attempt of performing unsupervised pre-training for improving the generalization ability of the lea…

Data AugmentationPerson Re-IdentificationUnsupervised Pre-training

Domain Adaptive Egocentric Person Re-identification

2021-03-08 · Ankit Choudhary, Deepak Mishra, Arnab Karmakar

Person re-identification (re-ID) in first-person (egocentric) vision is a fairly new and unexplored problem. With the increase of wearable video recording devices, egocentric data becomes readily available, and person re…

Person Re-IdentificationStyle Transfer