paper-with-me

홈 › Papers

Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos

2026-06-30 · Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado López, Mathias Unberath arxiv

Vision foundation models such as SAM 3 can provide transferable object-level structure across diverse surgical video conditions, but segmentation outputs do not explicitly encode the action-conditioned semantics that define functional surgical landmarks. Estimating instrument extent and geometry differs from localizing the tip or anchor relevant to clipping, grasping, or dissecting. We investigate vision foundation model-enabled sparse action-aware landmark localization, using zero-shot, point-prompted structural masks to provide dense instrument-level context without manual pixel-level mask annotations. We propose a lightweight refinement framework that uses SAM 3 as a structural prior. A coarse multi-frame network predicts tip and anchor prompts, generating non-oracle masks that are fused with visual and heatmap features to refine functional landmark predictions. We compare direct mask-augmented supervision, prediction-derived mask-prior refinement, and auxiliary mask supervision to examine how vision foundation model-derived structure should enter a precision-oriented localization system. Experiments on 7,867 clips from 60 surgical videos spanning YouTube, Cholec80, HeiChole, SurgVU, and CRCD evaluate the approach under heterogeneous conditions. Without manual pixel-level mask annotations for training, the proposed model achieves overall F1 scores of 72.4% for tip and 58.0% for anchor localization. Ablations show that coarse-to-fine refinement provides a substantial performance gain, while prediction-derived structural priors provide additional improvement when incorporated as intermediate guidance rather than direct localization targets.

📄 PDF Abstract BibTeX arXiv:2606.31007

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

FreeEnricher: Enriching Face Landmarks without Additional Cost

2022-12-19 · Yangyu Huang, Xi Chen, Jongyoo Kim, Hao Yang 외

Recent years have witnessed significant growth of face alignment. Though dense facial landmark is highly demanded in various scenarios, e.g., cosmetic medicine and facial beautification, most works only consider sparse f…

Face Alignment

Unsupervised Landmark Detection Based Spatiotemporal Motion Estimation for 4D Dynamic Medical Images

2021-09-30 · Yuyu Guo, Lei Bi, Dongming Wei, Liyun Chen 외

Motion estimation is a fundamental step in dynamic medical image processing for the assessment of target organ anatomy and function. However, existing image-based motion estimation methods, which optimize the motion fiel…

AnatomyMotion EstimationUnsupervised Landmark Detection

RaCalNet: Radar Calibration Network for Sparse-Supervised Metric Depth Estimation

2025-06-18 · Xingrui Qin, Wentao Zhao, Chuan Cao, Yihe Niu 외

Dense metric depth estimation using millimeter-wave radar typically requires dense LiDAR supervision, generated via multi-frame projection and interpolation, to guide the learning of accurate depth from sparse radar meas…

Depth EstimationDepth Prediction

SparseFormer: Attention-based Depth Completion Network

2022-06-09 · Frederik Warburg, Michael Ramamonjisoa, Manuel López-Antequera

Most pipelines for Augmented and Virtual Reality estimate the ego-motion of the camera by creating a map of sparse 3D landmarks. In this paper, we tackle the problem of depth completion, that is, densifying this sparse 3…

Depth Completion

3D face reconstruction with dense landmarks

2022-04-06 · Erroll Wood, Tadas Baltrusaitis, Charlie Hewitt, Matthew Johnson 외

Landmarks often play a key role in face analysis, but many aspects of identity or expression cannot be represented by sparse landmarks alone. Thus, in order to reconstruct faces more accurately, landmarks are often combi…

3D Face ReconstructionCPUFace AlignmentFace Model+1