paper-with-me

Papers

Tracking Meets Large Multimodal Models for Driving Scenario Understanding

2025-03-18 · Ayesha Ishaq, Jean Lahoud, Fahad Shahbaz Khan, Salman Khan, Hisham Cholakkal, Rao Muhammad Anwer

Large Multimodal Models (LMMs) have recently gained prominence in autonomous driving research, showcasing promising capabilities across various emerging benchmarks. LMMs specifically designed for this domain have demonstrated effective perception, planning, and prediction skills. However, many of these methods underutilize 3D spatial and temporal elements, relying mainly on image data. As a result, their effectiveness in dynamic driving environments is limited. We propose to integrate tracking information as an additional input to recover 3D spatial and temporal details that are not effectively captured in the images. We introduce a novel approach for embedding this tracking information into LMMs to enhance their spatiotemporal understanding of driving scenarios. By incorporating 3D tracking data through a track encoder, we enrich visual queries with crucial spatial and temporal cues while avoiding the computational overhead associated with processing lengthy video sequences or extensive 3D inputs. Moreover, we employ a self-supervised approach to pretrain the tracking encoder to provide LMMs with additional contextual information, significantly improving their performance in perception, planning, and prediction tasks for autonomous driving. Experimental results demonstrate the effectiveness of our approach, with a gain of 9.5% in accuracy, an increase of 7.04 points in the ChatGPT score, and 9.4% increase in the overall score over baseline models on DriveLM-nuScenes benchmark, along with a 3.7% final score improvement on DriveLM-CARLA. Our code is available at https://github.com/mbzuai-oryx/TrackingMeetsLMM

📄 PDF Abstract BibTeX arXiv:2503.14498

Code (1)

mbzuai-oryx/trackingmeetslmm 공식 구현 pytorch

Tasks

Autonomous Driving

Similar Papers 제목 키워드 기반

Know Your Surroundings: Panoramic Multi-Object Tracking by Multimodality Collaboration

2021-05-31 · Yuhang He, Wentao Yu, Jie Han, Xing Wei 외

In this paper, we focus on the multi-object tracking (MOT) problem of automatic driving and robot navigation. Most existing MOT methods track multiple objects using a singular RGB camera, which are prone to camera field-…

Multi-Object TrackingObject TrackingRobot Navigation

Underwater Camouflaged Object Tracking Meets Vision-Language SAM2

2024-09-25 · Chunhui Zhang, Li Liu, Guanjie Huang, Zhipeng Zhang 외

Over the past decade, significant progress has been made in visual object tracking, largely due to the availability of large-scale datasets. However, these datasets have primarily focused on open-air scenarios and have l…

ObjectObject TrackingVideo SegmentationVideo Semantic Segmentation+1

Drive-JEPA: Video JEPA Meets Multimodal Trajectory Distillation for End-to-End Driving

2026-01-29 · Linhan Wang, Zichong Yang, Chen Bai, Guoxiang Zhang 외 arxiv

End-to-end autonomous driving increasingly leverages self-supervised video pretraining to learn transferable planning representations. However, pretraining video world models for scene understanding has so far brought on…

Scene UnderstandingTrajectory PlanningAutonomous Driving

VideoMolmo: Spatio-Temporal Grounding Meets Pointing

2025-06-05 · Ghazi Shazan Ahmad, Ahmed Heakl, Hanan Gani, Abdelrahman Shaker 외

Spatio-temporal localization is vital for precise interactions across diverse domains, from biological research to autonomous navigation and interactive interfaces. Current video-based approaches, while proficient in tra…

Autonomous DrivingAutonomous NavigationCell TrackingReferring Video Object Segmentation+4

There is More than Meets the Eye: Self-Supervised Multi-Object Detection and Tracking with Sound by Distilling Multimodal Knowledge

2021-03-01 · CVPR 2021 1 · Francisco Rivera Valverde, Juana Valeria Hurtado, Abhinav Valada

Attributes of sound inherent to objects can provide valuable cues to learn rich representations for object detection and tracking. Furthermore, the co-occurrence of audiovisual events in videos can be exploited to locali…

object-detectionObject Detection