MPPNet: Multi-Frame Feature Intertwining with Proxy Points for 3D Temporal Object Detection
Accurate and reliable 3D detection is vital for many applications including autonomous driving vehicles and service robots. In this paper, we present a flexible and high-performance 3D detection framework, named MPPNet, for 3D temporal object detection with point cloud sequences. We propose a novel three-hierarchy framework with proxy points for multi-frame feature encoding and interactions to achieve better detection. The three hierarchies conduct per-frame feature encoding, short-clip feature fusion, and whole-sequence feature aggregation, respectively. To enable processing long-sequence point clouds with reasonable computational resources, intra-group feature mixing and inter-group feature attention are proposed to form the second and third feature encoding hierarchies, which are recurrently applied for aggregating multi-frame trajectory features. The proxy points not only act as consistent object representations for each frame, but also serve as the courier to facilitate feature interaction between frames. The experiments on large Waymo Open dataset show that our approach outperforms state-of-the-art methods with large margins when applied to both short (e.g., 4-frame) and long (e.g., 16-frame) point cloud sequences. Code is available at https://github.com/open-mmlab/OpenPCDet.
Code (1)
Tasks
Autonomous Drivingobject-detectionObject DetectionMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Care for the Mind Amid Chronic Diseases: An Interpretable AI Approach Using IoT
Health sensing for chronic disease management creates immense benefits for social welfare. Existing health sensing studies primarily focus on the prediction of physical chronic diseases. Depression, a widespread complica…
Decision MakingDepression DetectionManagementPredictionMODfinity: Unsupervised Domain Adaptation with Multimodal Information Flow Intertwining
Multimodal unsupervised domain adaptation leverages unlabeled data in the target domain to enhance multimodal systems continuously. While current state-of-the-art methods encourage interaction between sub-models of d…
Domain AdaptationModel Selectionmultimodal interactionUnsupervised Domain AdaptationMulti-Scale Context Intertwining for Semantic Segmentation
Accurate semantic image segmentation requires the joint consideration of local appearance, semantic information, and global scene context. In todayâs age of pre-trained deep networks and their powerful convolutional fe…
Image SegmentationSegmentationSemantic SegmentationFrame Fusion with Vehicle Motion Prediction for 3D Object Detection
In LiDAR-based 3D detection, history point clouds contain rich temporal information helpful for future prediction. In the same way, history detections should contribute to future detections. In this paper, we propose a d…
3D Object DetectionFuture predictionModel Selectionmotion prediction+2Conceptual Modeling and Artificial Intelligence: A Systematic Mapping Study
In conceptual modeling (CM), humans apply abstraction to represent excerpts of reality for means of understanding and communication, and processing by machines. Artificial Intelligence (AI) is applied to vast amounts of …