paper-with-me

Papers

TGRIP: A Text-Guided Approach to Vehicle Instance Prediction in Autonomous Driving

2026-07-06 · Miguel Antunes-García, Santiago Montiel-Marín, Fabio Sánchez-García, Rodrigo Gutiérrez-Moreno, Rafael Barea, Luis M. Bergasa arxiv

Bird's-Eye View (BEV) end-to-end instance prediction has emerged as a robust paradigm for autonomous driving perception, effectively mitigating the error propagation inherent in traditional modular pipelines. However, current state-of-the-art approaches rely predominantly on geometric supervision, such as occupancy regression and optical flow, effectively treating scene agents as generic moving obstacles. This absence of explicit semantic awareness imposes limitations on the capacity of the model to solve ambiguities in complex scenarios, particularly those where object-specific behavior is essential for accurate forecasting (e.g. overtaking, intersections). In this paper, we introduce Text-Guided Representation for Instance Prediction (TGRIP), a novel framework that bridges this gap by injecting rich semantic priors into the instance prediction loop. The proposed teacher-student pipeline employs Vision-Language Foundation Models to generate dense, semantic-enhanced BEV maps from multi-camera images. These maps serve as auxiliary supervision during training, guiding the network to learn spatio-temporal representations that are not only geometrically consistent but also semantically discriminative. To the best of our knowledge, this represents the first attempt to unify semantic guidance with the temporal task of future instance prediction. The experimental results demonstrate that TGRIP surpasses existing state-of-the-art models in nuScenes, validating the hypothesis that semantic enrichment is a fundamental element for robust, end-to-end motion prediction. Code is available on https://github.com/miguelag99/TGRIP.

📄 PDF Abstract BibTeX arXiv:2607.04812

Code (0)

등록된 구현이 없습니다.

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

Part-Guided Attention Learning for Vehicle Instance Retrieval

2019-09-13 · Xin-Yu Zhang, Rufeng Zhang, Jiewei Cao, Dong Gong 외

Vehicle instance retrieval often requires one to recognize the fine-grained visual differences between vehicles. Besides the holistic appearance of vehicles which is easily affected by the viewpoint variation and distort…

Fine-Grained Image ClassificationRetrievalVehicle Re-Identification

Vehicle Behavior Prediction by Episodic-Memory Implanted NDT

2024-02-13 · Peining Shen, Jianwu Fang, Hongkai Yu, Jianru Xue

In autonomous driving, predicting the behavior (turning left, stopping, etc.) of target vehicles is crucial for the self-driving vehicle to make safe decisions and avoid accidents. Existing deep learning-based methods ha…

Autonomous DrivingPrediction

Curriculum Learning in Genetic Programming Guided Local Search for Large-scale Vehicle Routing Problems

2025-05-17 · Saining Liu, Yi Mei, Mengjie Zhang

Manually designing (meta-)heuristics for the Vehicle Routing Problem (VRP) is a challenging task that requires significant domain expertise. Recently, data-driven approaches have emerged as a promising solution, automati…

TrafficPredict: Trajectory Prediction for Heterogeneous Traffic-Agents

2018-11-06 · Yuexin Ma, Xinge Zhu, Sibo Zhang, Ruigang Yang 외

To safely and efficiently navigate in complex urban traffic, autonomous vehicles must make responsible predictions in relation to surrounding traffic-agents (vehicles, bicycles, pedestrians, etc.). A challenging and crit…

Autonomous VehiclesNavigatePredictionTraffic Prediction+1

Count Anything

2026-05-29 · Mengqi Lei, Shuokun Cheng, Wei Bao, Shaoyi Du 외 arxiv

Object counting remains fragmented across domain-specific datasets and task formulations, despite rapid progress in generalist vision models. Existing counting models are often tailored to scenarios such as crowds, vehic…

Domain GeneralizationObject Counting