paper-with-me

홈 › Papers

RepSAM: Bridging Foundation Models to Robotic Vision via Representation-Guided Adaptation

2026-05-25 · Wenhui Chu arxiv

Robotic perception in unstructured environments remains challenging despite the zero-shot capabilities of foundation models such as SAM. This work attributes performance degradation to non-uniform representation shifts across transformer layers: shallow layers exhibit substantial domain gaps (CKA < 0.5), whereas deep layers transfer effectively (CKA > 0.7). Based on this observation, we propose RepSAM, a representation-guided parameter-efficient fine-tuning (PEFT) framework for adapting foundation models to robotic vision. RepSAM employs a theoretically grounded CKA-guided rank allocation strategy combined with a multi-modal fusion module for robust handling of challenging robotic scenarios, including transparent objects and cluttered scenes. Experimental evaluation across six benchmarks and robotic manipulation tasks demonstrates that RepSAM achieves 97.9% of full fine-tuning performance (89.0% vs. 90.9% mIoU) while reducing trainable parameters by 158x (from 632M to 4.0M). RepSAM outperforms DoRA by 7.9% mIoU with just 4 hours of training on a single A100 GPU (a 96x reduction from full fine-tuning, which takes 384 GPU-hours). These improvements are statistically significant (p < 0.01) and translate to a 12.0% absolute improvement in robotic manipulation success rates over the LoRA (RGB) baseline.

📄 PDF Abstract BibTeX arXiv:2605.25495

Code (0)

등록된 구현이 없습니다.

Tasks

parameter-efficient fine-tuning

Similar Papers 제목 키워드 기반

AsyncVLA: An Asynchronous VLA for Fast and Robust Navigation on the Edge

2026-02-13 · Noriaki Hirose, Catherine Glossop, Dhruv Shah, Sergey Levine arxiv

Robotic foundation models achieve strong generalization by leveraging internet-scale vision-language representations, but their massive computational cost creates a fundamental bottleneck: high inference latency. In dyna…

Learning Structured Robot Policies from Vision-Language Models via Synthetic Neuro-Symbolic Supervision

2026-04-03 · Alessandro Adami, Tommaso Tubaldo, Marco Todescato, Ruggero Carli 외 arxiv

Vision-Language Models (VLMs) have recently demonstrated strong capabilities in mapping multimodal observations to robot behaviors. However, most current approaches rely on end-to-end visuomotor policies that remain opaq…

Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy

2025-11-14 · Vinit Mehta, Charu Sharma, Karthick Thiyagarajan arxiv

With the rapid advancement of artificial intelligence and robotics, the integration of Large Language Models (LLMs) with 3D vision is emerging as a transformative approach to enhancing robotic sensing technologies. This …

Scene Understanding3D Generation

Foundation Model Driven Robotics: A Comprehensive Review

2025-07-14 · Muhammad Tayyab Khan, Ammar Waheed arxiv

The rapid emergence of foundation models, particularly Large Language Models (LLMs) and Vision-Language Models (VLMs), has introduced a transformative paradigm in robotics. These models offer powerful capabilities in sem…

Multimodal ReasoningScene Generation

SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding

2025-11-21 · Nikolay Nikolov, Giuliano Albanese, Sombit Dey, Aleksandar Yanev 외 arxiv

Robotic Foundation Models (RFMs) hold great promise as generalist, end-to-end systems for robot control. Yet their ability to generalize across new environments, tasks, and embodiments remains limited. We argue that a ma…

Spatial Reasoning