paper-with-me

홈 › Papers

RoPETR: Improving Temporal Camera-Only 3D Detection by Integrating Enhanced Rotary Position Embedding

2025-04-17 · Hang Ji, Tao Ni, Xufeng Huang, Tao Luo, Xin Zhan, Junbo Chen

This technical report introduces a targeted improvement to the StreamPETR framework, specifically aimed at enhancing velocity estimation, a critical factor influencing the overall NuScenes Detection Score. While StreamPETR exhibits strong 3D bounding box detection performance as reflected by its high mean Average Precision our analysis identified velocity estimation as a substantial bottleneck when evaluated on the NuScenes dataset. To overcome this limitation, we propose a customized positional embedding strategy tailored to enhance temporal modeling capabilities. Experimental evaluations conducted on the NuScenes test set demonstrate that our improved approach achieves a state-of-the-art NDS of 70.86% using the ViT-L backbone, setting a new benchmark for camera-only 3D object detection.

📄 PDF Abstract BibTeX arXiv:2504.12643

Code (0)

등록된 구현이 없습니다.

Tasks

3D Object Detectionobject-detectionObject DetectionPosition

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically

Similar Papers 제목 키워드 기반

Capturing Temporal Dynamics in Large-Scale Canopy Tree Height Estimation

2025-01-31 · Jan Pauls, Max Zimmer, Berkant Turan, Sassan Saatchi 외

With the rise in global greenhouse gas emissions, accurate large-scale tree canopy height maps are essential for understanding forest structure, estimating above-ground biomass, and monitoring ecological disruptions. To …

BEVFusion4D: Learning LiDAR-Camera Fusion Under Bird's-Eye-View via Cross-Modality Guidance and Temporal Aggregation

2023-03-30 · Hongxiang Cai, Zeyuan Zhang, Zhenyu Zhou, Ziyin Li 외

Integrating LiDAR and Camera information into Bird's-Eye-View (BEV) has become an essential topic for 3D object detection in autonomous driving. Existing methods mostly adopt an independent dual-branch framework to gener…

3D Object DetectionAutonomous Drivingobject-detectionObject Detection

SFOD: Spiking Fusion Object Detector

2024-03-22 · CVPR 2024 1 · Yimeng Fan, Wei zhang, Changsong Liu, Mingyang Li 외

Event cameras, characterized by high temporal resolution, high dynamic range, low power consumption, and high pixel bandwidth, offer unique capabilities for object detection in specialized contexts. Despite these advanta…

Objectobject-detectionObject Detection

Graph and Temporal Convolutional Networks for 3D Multi-person Pose Estimation in Monocular Videos

2020-12-22 · Yu Cheng, Bo wang, Bo Yang, Robby T. Tan

Despite the recent progress, 3D multi-person pose estimation from monocular videos is still challenging due to the commonly encountered problem of missing information caused by occlusion, partially out-of-frame target pe…

3D Absolute Human Pose Estimation3D Human Pose Estimation3D Multi-Person Pose Estimation3D Multi-Person Pose Estimation (absolute)+6

DenseBEV: Transforming BEV Grid Cells into 3D Objects

2025-12-18 · Marius Dähling, Sebastian Krebs, J. Marius Zöllner arxiv

In current research, Bird's-Eye-View (BEV)-based transformers are increasingly utilized for multi-camera 3D object detection. Traditional models often employ random queries as anchors, optimizing them successively. Recen…

Pedestrian Detection3D Object Detection