paper-with-me

홈 › Papers

EndoVLA: Dual-Phase Vision-Language-Action Model for Autonomous Tracking in Endoscopy

2025-05-21 · Chi Kit Ng, Long Bai, Guankun Wang, Yupeng Wang, Huxin Gao, Kun Yuan, Chenhan Jin, Tieyong Zeng, Hongliang Ren

In endoscopic procedures, autonomous tracking of abnormal regions and following circumferential cutting markers can significantly reduce the cognitive burden on endoscopists. However, conventional model-based pipelines are fragile for each component (e.g., detection, motion planning) requires manual tuning and struggles to incorporate high-level endoscopic intent, leading to poor generalization across diverse scenes. Vision-Language-Action (VLA) models, which integrate visual perception, language grounding, and motion planning within an end-to-end framework, offer a promising alternative by semantically adapting to surgeon prompts without manual recalibration. Despite their potential, applying VLA models to robotic endoscopy presents unique challenges due to the complex and dynamic anatomical environments of the gastrointestinal (GI) tract. To address this, we introduce EndoVLA, designed specifically for continuum robots in GI interventions. Given endoscopic images and surgeon-issued tracking prompts, EndoVLA performs three core tasks: (1) polyp tracking, (2) delineation and following of abnormal mucosal regions, and (3) adherence to circular markers during circumferential cutting. To tackle data scarcity and domain shifts, we propose a dual-phase strategy comprising supervised fine-tuning on our EndoVLA-Motion dataset and reinforcement fine-tuning with task-aware rewards. Our approach significantly improves tracking performance in endoscopy and enables zero-shot generalization in diverse scenes and complex sequential tasks.

📄 PDF Abstract BibTeX arXiv:2505.15206

Code (0)

등록된 구현이 없습니다.

Tasks

Motion PlanningVision-Language-ActionZero-shot Generalization

Similar Papers 제목 키워드 기반

TMR-VLA:Vision-Language-Action Model for Magnetic Motion Control of Tri-leg Silicone-based Soft Robot

2026-02-28 · Ruijie Tang, Chi Kit Ng, Kaixuan Wu, Long Bai 외 arxiv

In-vivo environments, magnetically actuated soft robots offer advantages such as wireless operation and precise control, showing promising potential for painless detection and therapeutic procedures. We developed a trile…

ReSW-VL: Representation Learning for Surgical Workflow Analysis Using Vision-Language Model

2025-05-19 · Satoshi Kondo

Surgical phase recognition from video is a technology that automatically classifies the progress of a surgical procedure and has a wide range of potential applications, including real-time surgical support, optimization …

Language ModelingLanguage ModellingPrompt LearningRepresentation Learning+2

PhaForce: Phase-Scheduled Visual-Force Policy Learning with Slow Planning and Fast Correction for Contact-Rich Manipulation

2026-03-09 · Mingxin Wang, Zhirun Yue, Renhao Lu, Yizhe Li 외 arxiv

Contact-rich manipulation requires not only vision-dominant task semantics but also closed-loop reactions to force/torque (F/T) transients. Yet, generative visuomotor policies are typically constrained to low-frequency u…

PAMAE: Phase-Aware-MoE Action Experts Towards Reliable Flow-Matching Vision-Language-Action Policies

2026-06-25 · Jiayu Yang, Tao Yang, Xiang Chang, Fei Chao 외 arxiv

Reliable action generation for multi-stage robotic manipulation remains challenging for Vision-Language-Action (VLA) models. While existing flow-matching VLA policies offer strong multimodal grounding and generalization,…

Move-Then-Operate: Behavioral Phasing for Human-Like Robotic Manipulation

2026-04-26 · Haoming Xu, Lei Lei, Jie Gu, Chu Tang 외 arxiv

We present Move-Then-Operate, a Vision language action framework that explicitly decouples robotic manipulation into two distinct behavioral phases: coarse relocation (move) and contact-critical interaction (operate). Un…