paper-with-me

Papers

SecondPose: SE(3)-Consistent Dual-Stream Feature Fusion for Category-Level Pose Estimation

2023-11-18 · CVPR 2024 1 · Yamei Chen, Yan Di, Guangyao Zhai, Fabian Manhardt, Chenyangguang Zhang, Ruida Zhang, Federico Tombari, Nassir Navab, Benjamin Busam

Category-level object pose estimation, aiming to predict the 6D pose and 3D size of objects from known categories, typically struggles with large intra-class shape variation. Existing works utilizing mean shapes often fall short of capturing this variation. To address this issue, we present SecondPose, a novel approach integrating object-specific geometric features with semantic category priors from DINOv2. Leveraging the advantage of DINOv2 in providing SE(3)-consistent semantic features, we hierarchically extract two types of SE(3)-invariant geometric features to further encapsulate local-to-global object-specific information. These geometric features are then point-aligned with DINOv2 features to establish a consistent object representation under SE(3) transformations, facilitating the mapping from camera space to the pre-defined canonical space, thus further enhancing pose estimation. Extensive experiments on NOCS-REAL275 demonstrate that SecondPose achieves a 12.4% leap forward over the state-of-the-art. Moreover, on a more complex dataset HouseCat6D which provides photometrically challenging objects, SecondPose still surpasses other competitors by a large margin.

📄 PDF Abstract BibTeX arXiv:2311.11125

Code (1)

NOrangeeroli/SecondPose 공식 구현 pytorch

Tasks

ObjectPose Estimation

Similar Papers 제목 키워드 기반

LI-DSN: A Layer-wise Interactive Dual-Stream Network for EEG Decoding

2026-04-02 · Chenghao Yue, Zhiyuan Ma, Zhongye Xia, Xinche Zhang 외 arxiv

Electroencephalography (EEG) provides a non-invasive window into brain activity, offering high temporal resolution crucial for understanding and interacting with neural processes through brain-computer interfaces (BCIs).…

Emotion RecognitionEeg Decoding

physfusion: A Transformer-based Dual-Stream Radar and Vision Fusion Framework for Open Water Surface Object Detection

2026-03-02 · Yuting Wan, Liguo Sun, Jiuwu Hao, Zao Zhang 외 arxiv

Detecting water-surface targets for Unmanned Surface Vehicles (USVs) is challenging due to wave clutter, specular reflections, and weak appearance cues in long-range observations. Although 4D millimeter-wave radar comple…

Object DetectionPoint Clouds

P2Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation

2026-08-13 · Yi Shi, Huichao Xie, Yuqing Wang, Mingyu Wang 외 arxiv

Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundamental challenge. Existing prior-guided m…

Object Detection

DualDiff: Dual-branch Diffusion Model for Autonomous Driving with Semantic Fusion

2025-05-03 · Haoteng Li, Zhao Yang, Zezhong Qian, Gongpeng Zhao 외

Accurate and high-fidelity driving scene reconstruction relies on fully leveraging scene information as conditioning. However, existing approaches, which primarily use 3D bounding boxes and binary maps for foreground and…

3D Object DetectionAutonomous DrivingBEV Segmentationobject-detection+2

DCMorph: Face Morphing via Dual-Stream Cross-Attention Diffusion

2026-04-23 · Tahar Chettaoui, Eduarda Caldeira, Guray Ozgur, Raghavendra Ramachandra 외 arxiv

Advancing face morphing attack techniques is crucial to anticipate evolving threats and develop robust defensive mechanisms for identity verification systems. This work introduces DCMorph, a dual-stream diffusion-based m…

Face Recognition