paper-with-me

홈 › Papers

Mind the Privileged-to-Camera Gap: Actor-Centric Sidecar Supervision for Camera-First Open-Loop Waypoint Prediction

2026-06-18 · Feeza Khan Khanzada, Jaerock Kwon arxiv

Camera-first autonomous-driving models predict future ego waypoints from images, ego-state features, and route commands, but waypoint supervision alone does not explicitly supervise actor-level representations of nearby road users. We study this as supervised representation learning for open-loop waypoint prediction. The deployable model uses multi-view RGB, ego state, and route command at inference. During training, simulator-derived sidecar labels supervise actor grounding, privileged hindsight actor relevance relative to the logged ego trajectory, and selected-actor short-horizon motion; these labels are never inference inputs. We evaluate route-disjoint splits with matched architecture, optimizer, validation criterion, checkpoint selection, and three seeds. A plain waypoint-only RGB baseline obtains 1.815$\pm$0.02 m final displacement error (FDE), and the matched no-teacher non-sidecar RGB control obtains 1.716$\pm$0.02 m. Road-user sidecar supervision (RU-sidecar) reduces FDE to 1.223$\pm$0.01 m, a 32.6% reduction over the plain baseline and 28.7% over the matched no-teacher non-sidecar RGB control. It improves over the plain baseline on 1445/1494 routes and over the matched no-teacher non-sidecar RGB control on 1417/1494 routes. Actor-conditioned slices show gains in all nonempty subsets, including 29.1% reduction for samples with at least four valid sidecar actors and 30.0% when a vulnerable road user is present. Optional simulator-state teacher alignment reaches 1.186$\pm$0.15 m FDE, but higher seed variability makes it secondary. Non-deployable simulator-state diagnostics remain stronger, indicating a privileged-to-camera gap. The evidence is limited to open-loop simulation diagnostics.

📄 PDF Abstract BibTeX arXiv:2606.20772

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Pictura: Perspective-View Self-Play at Scale for Driving

2026-07-28 · Yuan Yin, Elias Ramzi, Marc Lafon, Valentin Charraut 외 arxiv

Self-play in simulation produces robust driving policies at scale. Demonstrations of such behavior have been made using privileged vectorized observations such as exact poses and velocities, even for occluded agents. Thi…

Learning Ego-Centric BEV Representations from a Perspective-Privileged View: Cross-View Supervision for Online HD Map Construction

2026-05-12 · Daniel Lengerer, Mathias Pechinger, Klaus Bogenberger, Carsten Markgraf arxiv

Bird's-eye-view (BEV) representations derived from multi-camera input have become a central interface for online high-definition (HD) map construction. However, most approaches rely solely on ego-centric supervision, req…

Representation Learning

EgoToM: Benchmarking Theory of Mind Reasoning from Egocentric Videos

2025-03-28 · YuXuan Li, Vijay Veerabadran, Michael L. Iuzzolino, Brett D. Roads 외

We introduce EgoToM, a new video question-answering benchmark that extends Theory-of-Mind (ToM) evaluation to egocentric domains. Using a causal ToM model, we generate multi-choice video QA instances for the Ego4D datase…

BenchmarkingQuestion AnsweringVideo Question Answering

Symbiotic Attention with Privileged Information for Egocentric Action Recognition

2020-02-08 · Xiaohan Wang, Yu Wu, Linchao Zhu, Yi Yang

Egocentric video recognition is a natural testbed for diverse interaction reasoning. Due to the large action vocabulary in egocentric video datasets, recent studies usually utilize a two-branch structure for action recog…

Action RecognitionEgocentric Activity RecognitionGeneral Classificationobject-detection+3

Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind

2024-04-07 · Chiara Plizzari, Shubham Goel, Toby Perrett, Jacob Chalk 외

As humans move around, performing their daily tasks, they are able to recall where they have positioned objects in their environment, even if these objects are currently out of their sight. In this paper, we aim to mimic…