paper-with-me

홈 › Papers

StreamMOTP: Streaming and Unified Framework for Joint 3D Multi-Object Tracking and Trajectory Prediction

2024-06-28 · Jiaheng Zhuang, Guoan Wang, Siyu Zhang, Xiyang Wang, Hangning Zhou, Ziyao Xu, Chi Zhang, Zhiheng Li

3D multi-object tracking and trajectory prediction are two crucial modules in autonomous driving systems. Generally, the two tasks are handled separately in traditional paradigms and a few methods have started to explore modeling these two tasks in a joint manner recently. However, these approaches suffer from the limitations of single-frame training and inconsistent coordinate representations between tracking and prediction tasks. In this paper, we propose a streaming and unified framework for joint 3D Multi-Object Tracking and trajectory Prediction (StreamMOTP) to address the above challenges. Firstly, we construct the model in a streaming manner and exploit a memory bank to preserve and leverage the long-term latent features for tracked objects more effectively. Secondly, a relative spatio-temporal positional encoding strategy is introduced to bridge the gap of coordinate representations between the two tasks and maintain the pose-invariance for trajectory prediction. Thirdly, we further improve the quality and consistency of predicted trajectories with a dual-stream predictor. We conduct extensive experiments on popular nuSences dataset and the experimental results demonstrate the effectiveness and superiority of StreamMOTP, which outperforms previous methods significantly on both tasks. Furthermore, we also prove that the proposed framework has great potential and advantages in actual applications of autonomous driving.

📄 PDF Abstract BibTeX arXiv:2406.19844

Code (0)

등록된 구현이 없습니다.

Tasks

3D Multi-Object TrackingAutonomous DrivingMulti-Object TrackingObject TrackingPredictionTrajectory Prediction

Similar Papers 제목 키워드 기반

Uni-ASR: Unified LLM-Based Architecture for Non-Streaming and Streaming Automatic Speech Recognition

2026-03-11 · Yinfeng Xia, Jian Tang, Junfeng Hou, Gaopeng Xu 외 arxiv

Although the deep integration of the Automatic Speech Recognition (ASR) system with Large Language Models (LLMs) has significantly improved accuracy, the deployment of such systems in low-latency streaming scenarios rema…

Speech Recognition

StreamingEval: A Unified Evaluation Protocol towards Realistic Streaming Video Understanding

2026-03-23 · Guowei Tang, Tianwen Qian, Huanran Zheng, Yifei Wang 외 arxiv

Real-time, continuous understanding of visual signals is essential for real-world interactive AI applications, and poses a fundamental system-level challenge. Existing research on streaming video understanding, however, …

ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding

2026-01-15 · Xueyun Tian, Wei Li, Bingbing Xu, Heng Dong 외 arxiv

Recent Omni-multimodal Large Language Models show promise in unified audio, vision, and text modeling. However, streaming audio-video understanding remains challenging, as existing approaches suffer from disjointed capab…

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars

2026-07-25 · Quanyue Song, Yishan He, Yanbo Ding, Zhixiang He 외 arxiv

Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing a promising foundation for interactive avatar systems. However, ext…

Dual-mode ASR: Unify and Improve Streaming ASR with Full-context Modeling

2020-10-12 · ICLR 2021 1 · Jiahui Yu, Wei Han, Anmol Gulati, Chung-Cheng Chiu 외

Streaming automatic speech recognition (ASR) aims to emit each hypothesized word as quickly and accurately as possible, while full-context ASR waits for the completion of a full speech utterance before emitting completed…

Automatic Speech RecognitionAutomatic Speech Recognition (ASR)Knowledge Distillationspeech-recognition+1