paper-with-me

Papers

ForeSight: Multi-View Streaming Joint Object Detection and Trajectory Forecasting

2025-08-09 · Sandro Papais, Letian Wang, Brian Cheong, Steven L. Waslander arxiv

We introduce ForeSight, a novel joint detection and forecasting framework for vision-based 3D perception in autonomous vehicles. Traditional approaches treat detection and forecasting as separate sequential tasks, limiting their ability to leverage temporal cues. ForeSight addresses this limitation with a multi-task streaming and bidirectional learning approach, allowing detection and forecasting to share query memory and propagate information seamlessly. The forecast-aware detection transformer enhances spatial reasoning by integrating trajectory predictions from a multiple hypothesis forecast memory queue, while the streaming forecast transformer improves temporal consistency using past forecasts and refined detections. Unlike tracking-based methods, ForeSight eliminates the need for explicit object association, reducing error propagation with a tracking-free model that efficiently scales across multi-frame sequences. Experiments on the nuScenes dataset show that ForeSight achieves state-of-the-art performance, achieving an EPA of 54.9%, surpassing previous methods by 9.3%, while also attaining the best mAP and minADE among multi-view detection and forecasting models.

📄 PDF Abstract BibTeX arXiv:2508.07089

Code (0)

등록된 구현이 없습니다.

Tasks

Trajectory ForecastingAutonomous VehiclesSpatial ReasoningObject Detection

Results from the Paper

RankTaskDatasetModelMetrics
#35 Trajectory Prediction nuScenes ForeSight MinADE_10: 9.3

Similar Papers 제목 키워드 기반

Merlin:Empowering Multimodal LLMs with Foresight Minds

2023-11-30 · En Yu, Liang Zhao, Yana Wei, Jinrong Yang 외

Humans possess the remarkable ability to foresee the future to a certain extent based on present observations, a skill we term as foresight minds. However, this capability remains largely under explored within existing M…

Visual Question Answering

Experience-Embedded Visual Foresight

2019-11-12 · Lin Yen-Chen, Maria Bauza, Phillip Isola

Visual foresight gives an agent a window into the future, which it can use to anticipate events before they happen and plan strategic behavior. Although impressive results have been achieved on video prediction in constr…

PredictionVideo Prediction

Tell Me What's Next: Textual Foresight for Generic UI Representations

2024-06-12 · Andrea Burns, Kate Saenko, Bryan A. Plummer

Mobile app user interfaces (UIs) are rich with action, text, structure, and image content that can be utilized to learn generic UI representations for tasks like automating user commands, summarizing content, and evaluat…

Representation Learning

AirForesight: Current-to-Future Spatial Map Imagination with Cross-Space Planning Consistency for UAV-VLN

2026-08-13 · Yutong Liu, Xiaojie Li, Mingzhu Xu, Jianlong Wu arxiv

Unmanned Aerial Vehicle Vision-Language Navigation (UAV-VLN) requires agents to follow language instructions, infer spatial structure from sparse multi-view observations, and execute feasible 3D motion in complex outdoor…

Vision-Language NavigationTrajectory PredictionSpatial Reasoning

CLaD: Planning with Grounded Foresight via Cross-Modal Latent Dynamics

2026-03-31 · Andrew Jeong, Jaemin Kim, Sebin Lee, Sung-Eui Yoon arxiv

Robotic manipulation involves kinematic and semantic transitions that are inherently coupled via underlying actions. However, existing approaches plan within either semantic or latent space without explicitly aligning th…