paper-with-me

Papers

Cross-Modal RGB-D Fusion Transformer for 6D Pose Estimation of Non-Cooperative Spacecraft with Stereo-Derived Depth

2026-05-09 · Yongliang Zhen, Bo LÜ, Hang Yang, Xiaotian WU arxiv

On-orbit servicing and active debris removal involving non-cooperative spacecraft require reliable pose estimation to supply accurate position and orientation data for autonomous visual navigation. Learning-based monocular methods have seen widespread adoption in spacecraft pose estimation, yet they suffer from an intrinsic depth ambiguity problem and tend to fail under the harsh illumination conditions routinely encountered in orbit. Active depth sensors could in principle address the geometric ambiguity, but their power and mass requirements make them poorly suited to most spacecraft platforms. This work addresses these issues through a passive stereo vision framework for six-degree-of-freedom (6-DOF) pose estimation of non-cooperative spacecraft. A binocular stereo matching network called TSCA-Stereo is developed to cope with weak-texture surfaces, specular highlights, and severe lighting variations typical of space imagery. A cross-modal fusion Transformer is introduced to combine RGB appearance information with stereo depth features in an adaptive manner, supporting reliable pose recovery. A synthetic binocular multimodal dataset is also built for the experiments, covering stereo disparity maps and 6-DOF pose annotations across a range of illumination scenarios, attitude configurations, and noise levels. Experimental results show that TSCA-Stereo outperforms the baseline across every evaluated metric on this space-specific dataset. The full pose estimation pipeline achieves a mean translation error of 0.0419 m and a mean orientation error of 0.8632° under varied imaging conditions, confirming that the passive stereo approach is both effective and resilient when operating under the demanding visual conditions of the space environment.

📄 PDF Abstract BibTeX arXiv:2605.08592

Code (0)

등록된 구현이 없습니다.

Tasks

6D Pose EstimationVisual Navigation

Similar Papers 제목 키워드 기반

TransFusionOdom: Interpretable Transformer-based LiDAR-Inertial Fusion Odometry Estimation

2023-04-16 · Leyuan Sun, Guanqun Ding, Yue Qiu, Yusuke Yoshiyasu 외

Multi-modal fusion of sensors is a commonly used approach to enhance the performance of odometry estimation, which is also a fundamental module for mobile robots. However, the question of \textit{how to perform fusion am…

Sensor Fusion

Deep Fusion Transformer Network with Weighted Vector-Wise Keypoints Voting for Robust 6D Object Pose Estimation

2023-08-10 · ICCV 2023 1 · Jun Zhou, Kai Chen, Linlin Xu, Qi Dou 외

One critical challenge in 6D object pose estimation from a single RGBD image is efficient integration of two different modalities, i.e., color and depth. In this work, we tackle this problem by a novel Deep Fusion Transf…

6D Pose Estimation using RGBglobal-optimizationPose EstimationSemantic Similarity+1

UniCT Depth: Event-Image Fusion Based Monocular Depth Estimation with Convolution-Compensated ViT Dual SA Block

2025-07-26 · Luoxi Jing, Dianxi Shi, Zhe Liu, Songchang Jin 외 arxiv

Depth estimation plays a crucial role in 3D scene understanding and is extensively used in a wide range of vision tasks. Image-based methods struggle in challenging scenarios, while event cameras offer high dynamic range…

Monocular Depth EstimationScene Understanding

DAT: Dialogue-Aware Transformer with Modality-Group Fusion for Human Engagement Estimation

2024-10-11 · Jia Li, Yangchen Yu, Yin Chen, Yu Zhang 외

Engagement estimation plays a crucial role in understanding human social behaviors, attracting increasing research interests in fields such as affective computing and human-computer interaction. In this paper, we propose…

Cross-modal transformers for infrared and visible image fusion

2023-06-26 · IEEE Transactions on Circuits and Systems for Video Technology 2023 6 · Seonghyun Park, An Gia Vien, Chul Lee

Image fusion techniques aim to generate more informative images by merging multiple images of different modalities with complementary information. Despite significant fusion performance improvements of recent learning-ba…

Cross-Modal RetrievalDepth EstimationInfrared And Visible Image FusionMonocular Depth Estimation+2