paper-with-me

Papers

Deep Laparoscopic Stereo Matching with Transformers

2022-07-25 · Xuelian Cheng, Yiran Zhong, Mehrtash Harandi, Tom Drummond, Zhiyong Wang, ZongYuan Ge

The self-attention mechanism, successfully employed with the transformer structure is shown promise in many computer vision tasks including image recognition, and object detection. Despite the surge, the use of the transformer for the problem of stereo matching remains relatively unexplored. In this paper, we comprehensively investigate the use of the transformer for the problem of stereo matching, especially for laparoscopic videos, and propose a new hybrid deep stereo matching framework (HybridStereoNet) that combines the best of the CNN and the transformer in a unified design. To be specific, we investigate several ways to introduce transformers to volumetric stereo matching pipelines by analyzing the loss landscape of the designs and in-domain/cross-domain accuracy. Our analysis suggests that employing transformers for feature representation learning, while using CNNs for cost aggregation will lead to faster convergence, higher accuracy and better generalization than other options. Our extensive experiments on Sceneflow, SCARED2019 and dVPN datasets demonstrate the superior performance of our HybridStereoNet.

📄 PDF Abstract BibTeX arXiv:2207.12152

Code (1)

xueliancheng/hybridstereonet-main 공식 구현 pytorch

Tasks

object-detectionObject DetectionRepresentation LearningStereo Matching

Similar Papers 제목 키워드 기반

A Disparity Refinement Framework for Learning-based Stereo Matching Methods in Cross-domain Setting for Laparoscopic Images

2023-02-05 · Zixin Yang, Richard Simon, Cristian A. Linte

Purpose: Stereo matching methods that enable depth estimation are crucial for visualization enhancement applications in computer-assisted surgery (CAS). Learning-based stereo matching methods are promising to predict acc…

Depth EstimationStereo Matching

Automatic 3D Point Set Reconstruction from Stereo Laparoscopic Images using Deep Neural Networks

2016-07-31 · Balint Antal

In this paper, an automatic approach to predict 3D coordinates from stereo laparoscopic images is presented. The approach maps a vector of pixel intensities to 3D coordinates through training a six layer deep neural netw…

WT-MVSNet: Window-based Transformers for Multi-view Stereo

2022-05-28 · Jinli Liao, Yikang Ding, Yoli Shavit, Dihe Huang 외

Recently, Transformers were shown to enhance the performance of multi-view stereo by enabling long-range feature interaction. In this work, we propose Window-based Transformers (WT) for local feature matching and global …

Tracking Any Point Methods for Markerless 3D Tissue Tracking in Endoscopic Stereo Images

2025-08-11 · Konrad Reuter, Suresh Guttikonda, Sarah Latus, Lennart Maack 외 arxiv

Minimally invasive surgery presents challenges such as dynamic tissue motion and a limited field of view. Accurate tissue tracking has the potential to support surgical guidance, improve safety by helping avoid damage to…

Self-Supervised Depth Estimation in Laparoscopic Image using 3D Geometric Consistency

2022-08-17 · Baoru Huang, Jian-Qing Zheng, Anh Nguyen, Chi Xu 외

Depth estimation is a crucial step for image-guided intervention in robotic surgery and laparoscopic imaging system. Since per-pixel depth ground truth is difficult to acquire for laparoscopic image data, it is rarely po…

Depth Estimation