Rethinking Event-Based Object Detection through Representation-Level Temporal Aggregation and Model-Level Hypergraph Reasoning
Event cameras provide microsecond-level temporal resolution, low latency, and high dynamic range, offering potential for perception under fast motion and challenging illumination conditions. However, existing Event-based Object Detection (EOD) methods face limitations at both the representation and model levels: prior event representations usually encode temporal information indirectly through redundant structures, while detection models struggle to explicitly aggregate fragmented event responses into coherent high-order object features. To address these limitations, we present \textbf{Event Dual Temporal-Relational Aggregation Detector (Ev-DTAD)}, a unified EOD framework that integrates representation-level temporal encoding with model-level temporal-hypergraph reasoning. Specifically, we introduce \textbf{Hierarchical Temporal Aggregation (HTA)}, a compact three-channel pseudo-RGB representation that explicitly embeds temporal information across intra- and inter-window events. To further enhance detection under sparse and fragmented event responses, we propose \textbf{Frequency-aware Hypergraph Temporal Fusion (FHTF)}, which refines multi-scale event features through temporal evolution modeling and high-order relational reasoning. Extensive experiments on Gen1 (\textbf{+0.8 mAP}), 1Mpx/Gen4 (\textbf{+0.5 mAP}), and eTraM (\textbf{+3.0 mAP}) demonstrate that Ev-DTAD achieves a competitive accuracy--efficiency trade-off, validating the complementarity between compact temporal representation and temporal-hypergraph feature reasoning (Fig.~\ref{fig:bubble}). The code is available at: https://github.com/meisenwang/Ev-DTAD.
Code (0)
등록된 구현이 없습니다.
Tasks
Relational ReasoningObject DetectionSimilar Papers 제목 키워드 기반
A2DINOv3: Rethinking Multi-Modal Object Detection via Socialized Collaboration
Multi-modal object detection is essential for robust scene understanding in challenging conditions, including low-light and adverse environments. Recent vision foundation models (e.g., DINOv3) have exhibited strong repre…
Representation LearningScene UnderstandingAutonomous DrivingObject DetectionRemDet: Rethinking Efficient Model Design for UAV Object Detection
Object detection in Unmanned Aerial Vehicle (UAV) images has emerged as a focal area of research, which presents two significant challenges: i) objects are typically small and dense within vast images; ii) computational …
Objectobject-detectionObject DetectionSmall Object DetectionDetection Bank: An Object Detection Based Video Representation for Multimedia Event Recognition
While low-level image features have proven to be effective representations for visual recognition tasks such as object recognition and scene classification, they are inadequate to capture complex semantic meaning require…
Event DetectionObjectobject-detectionObject Detection+3Efficient Multi-Timescale Event Representations for Feed-Forward Object Detection
Autonomous systems require robust low-latency perception under rapidly changing scene dynamics and challenging illumination. In event cameras object detection commonly relies on recurrent architectures to accumulate spar…
Object DetectionDetZero: Rethinking Offboard 3D Object Detection with Long-term Sequential Point Clouds
Existing offboard 3D detectors always follow a modular pipeline design to take advantage of unlimited sequential point clouds. We have found that the full potential of offboard 3D detectors is not explored mainly due to …
3D Multi-Object Tracking3D Object DetectionObjectobject-detection+1