ShipTraj-R1: Reinforcing Ship Trajectory Prediction in Large Language Models via Group Relative Policy Optimization
Recent advancements in reinforcement fine-tuning have significantly improved the reasoning ability of large language models (LLMs). In particular, methods such as group relative policy optimization (GRPO) have demonstrated strong capabilities across various fields. However, applying LLMs to ship trajectory prediction remains largely unexplored. In this paper, we propose ShipTraj-R1, a novel LLM-based framework that reformulates ship trajectory prediction as a text-to-text generation problem. (1) We design a dynamic prompt containing trajectory information about conflicting ships to guide the model to achieve adaptive chain-of-thought (CoT) reasoning. (2) We introduce a comprehensive rule-based reward mechanism to incentivize the reasoning format and prediction accuracy of the model. (3) Our ShipTraj-R1 is reinforced through the GRPO mechanism guided by domain-specific prompts and rewards, and utilizes the Qwen3 as the model backbone. Extensive experimental results on two complex and real-world maritime datasets show that the proposed ShipTraj-R1 achieves the least error compared with state-of-the-art deep learning and LLM-based baselines.
Code (0)
등록된 구현이 없습니다.
Tasks
Trajectory PredictionText GenerationSimilar Papers 제목 키워드 기반
EnvShip-Bench: An Environment-Enhanced Benchmark for Short-Term Vessel Trajectory Prediction
Vessel trajectory prediction is important for intelligent shipping, maritime surveillance, and navigation safety. However, existing public maritime AIS resources are often limited by inconsistent forecasting protocols, u…
Trajectory ForecastingTrajectory PredictionForesight in Motion: Reinforcing Trajectory Prediction with Reward Heuristics
Motion forecasting for on-road traffic agents presents both a significant challenge and a critical necessity for ensuring safety in autonomous driving systems. In contrast to most existing data-driven approaches that dir…
Reinforcement LearningTrajectory PredictionAutonomous DrivingMotion ForecastingVideo Relation Detection with Trajectory-aware Multi-modal Features
Video relation detection problem refers to the detection of the relationship between different objects in videos, such as spatial relationship and action relationship. In this paper, we present video relation detection w…
Objectobject-detectionObject DetectionRelation+2Modeling Emergent Lexicon Formation with a Self-Reinforcing Stochastic Process
We introduce FiLex, a self-reinforcing stochastic process which models finite lexicons in emergent language experiments. The central property of FiLex is that it is a self-reinforcing process, parallel to the intuition t…
SocialMOIF: Multi-Order Intention Fusion for Pedestrian Trajectory Prediction
The analysis and prediction of agent trajectories are crucial for decision-making processes in intelligent systems, with precise short-term trajectory forecasting being highly significant across a range of applications. …
Pedestrian Trajectory PredictionTrajectory ForecastingTrajectory Prediction