paper-with-me

Papers

Data-centric Design of Learning-based Surgical Gaze Perception Models in Multi-Task Simulation

2026-02-09 · Yizhou Li, Shuyuan Yang, Jiaji Su, Zonghe Chua arxiv

In robot-assisted minimally invasive surgery (RMIS), reduced haptic feedback and depth cues increase reliance on expert visual perception, motivating gaze-guided training and learning-based surgical perception models. However, operative expert gaze is costly to collect, and it remains unclear how the source of gaze supervision, both expertise level (intermediate vs. novice) and perceptual modality (active execution vs. passive viewing), shapes what attention models learn. We introduce a paired active-passive, multi-task surgical gaze dataset collected on the da Vinci SimNow simulator across four drills. Active gaze was recorded during task execution using a VR headset with eye tracking, and the corresponding videos were reused as stimuli to collect passive gaze from observers, enabling controlled same-video comparisons. We quantify skill- and modality-dependent differences in gaze organization and evaluate the substitutability of passive gaze for operative supervision using fixation density overlap analyses and single-frame saliency modeling. Across settings, MSI-Net produced stable, interpretable predictions, whereas SalGAN was unstable and often poorly aligned with human fixations. Models trained on passive gaze recovered a substantial portion of intermediate active attention, but with predictable degradation, and transfer was asymmetric between active and passive targets. Notably, novice passive labels approximated intermediate-passive targets with limited loss on higher-quality demonstrations, suggesting a practical path for scalable, crowd-sourced gaze supervision in surgical coaching and perception modeling.

📄 PDF Abstract BibTeX arXiv:2602.09259

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding

2025-05-30 · Ege Özsoy, Arda Mamur, Felix Tristram, Chantal Pellegrini 외

Operating rooms (ORs) demand precise coordination among surgeons, nurses, and equipment in a fast-paced, occlusion-heavy environment, necessitating advanced perception models to enhance safety and efficiency. Existing da…

Action RecognitionGraph GenerationScene Graph Generation

EgoSurgery-Phase: A Dataset of Surgical Phase Recognition from Egocentric Open Surgery Videos

2024-05-30 · Ryo Fujii, Masashi Hatano, Hideo Saito, Hiroki Kajita

Surgical phase recognition has gained significant attention due to its potential to offer solutions to numerous demands of the modern operating room. However, most existing methods concentrate on minimally invasive surge…

Action RecognitionSurgical phase recognitionVideo Understanding

Unsupervised Gaze Prediction in Egocentric Videos by Energy-based Surprise Modeling

2020-01-30 · Sathyanarayanan N. Aakur, Arunkumar Bagavathi

Egocentric perception has grown rapidly with the advent of immersive computing devices. Human gaze prediction is an important problem in analyzing egocentric videos and has primarily been tackled through either saliency-…

Gaze PredictionPrediction

Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span

2025-11-23 · Heeseung Yun, Joonil Na, Jaeyeon Kim, Calvin Murdock 외 arxiv

People continuously perceive and interact with their surroundings based on underlying intentions that drive their exploration and behaviors. While research in egocentric user and scene understanding has focused primarily…

Scene Understanding

Human Gaze Guided Attention for Surgical Activity Recognition

2022-03-09 · Abdishakour Awale, Duygu Sarikaya

Modeling and automatically recognizing surgical activities are fundamental steps toward automation in surgery and play important roles in providing timely feedback to surgeons. Accurately recognizing surgical activities …

Activity RecognitionVideo Understanding