Multimodal Future Localization and Emergence Prediction for Objects in Egocentric View with a Reachability Prior
In this paper, we investigate the problem of anticipating future dynamics, particularly the future location of other vehicles and pedestrians, in the view of a moving vehicle. We approach two fundamental challenges: (1) the partial visibility due to the egocentric view with a single RGB camera and considerable field-of-view change due to the egomotion of the vehicle; (2) the multimodality of the distribution of future states. In contrast to many previous works, we do not assume structural knowledge from maps. We rather estimate a reachability prior for certain classes of objects from the semantic map of the present image and propagate it into the future using the planned egomotion. Experiments show that the reachability prior combined with multi-hypotheses learning improves multimodal prediction of the future location of tracked objects and, for the first time, the emergence of new objects. We also demonstrate promising zero-shot transfer to unseen datasets. Source code is available at $\href{https://github.com/lmb-freiburg/FLN-EPN-RPN}{\text{this https URL.}}$
Code (1)
Similar Papers 제목 키워드 기반
DiaLoc: An Iterative Approach to Embodied Dialog Localization
Multimodal learning has advanced the performance for many vision-language tasks. However, most existing works in embodied dialog research focus on navigation and leave the localization task understudied. The few existing…
Motion Transformer with Global Intention Localization and Local Movement Refinement
Predicting multimodal future behavior of traffic participants is essential for robotic vehicles to make safe decisions. Existing works explore to directly predict future trajectories based on latent features or utilize d…
motion predictionPredictionTrajectory PredictionEmergence of Structured Behaviors from Curiosity-Based Intrinsic Motivation
Infants are experts at playing, with an amazing ability to generate novel structured behaviors in unstructured environments that lack clear extrinsic reward signals. We seek to replicate some of these abilities with a ne…
motion predictionObjectMTR-A: 1st Place Solution for 2022 Waymo Open Dataset Challenge -- Motion Prediction
In this report, we present the 1st place solution for motion prediction track in 2022 Waymo Open Dataset Challenges. We propose a novel Motion Transformer framework for multimodal motion prediction, which introduces a sm…
motion predictionPredictionWhere Do Vision-Language Models Fail? World Scale Analysis for Image Geolocalization
Image geolocalization has traditionally been addressed through retrieval-based place recognition or geometry-based visual localization pipelines. Recent advances in Vision-Language Models (VLMs) have demonstrated strong …
Multimodal ReasoningVisual LocalizationImage Matching