paper-with-me

Papers

Ro-SLM: Onboard Small Language Models for Robot Task Planning and Operation Code Generation

2026-04-13 · Wenhao Wang, Yanyan Li, Long Jiao, Jiawei Yuan arxiv

Recent advances in large language models (LLMs) provide robots with contextual reasoning abilities to comprehend human instructions. Yet, current LLM-enabled robots typically depend on cloud-based models or high-performance computing infrastructure, which limit their deployment on robots under unreliable internet environments or with constrained computational resources, such as UAVs and small ground vehicles. Thus, deploying fine-tuned small language models (SLMs) that support onboard deployment offers a promising alternative. This paper introduces Ro-SLM, a framework that enables reliable SLM-driven robot operation by distilling LLMs' knowledge and reasoning. Ro-SLM starts from dataset synthesis by leveraging LLMs to generate diverse task instructions, produce corresponding ground truth code with minimal human assistance, and augment instructions into real-world application scenarios. Ro-SLM is then fine-tuned with the dataset, in which LLM serves as a reward function to guide the training. Extensive experiments on UAV operation tasks demonstrate that Ro-SLM improves the performance of SLM from being incapable of supporting robotic task planning and code generation to achieving performance that approaches LLM.

📄 PDF Abstract BibTeX arXiv:2604.10929

Code (0)

등록된 구현이 없습니다.

Tasks

Robot Task PlanningCode Generation

Similar Papers 제목 키워드 기반

Multi-Agent Robotic Control with Onboard Vision-Language Models

2026-07-08 · Kajetan Rachwał, Maciej Majek, Bartłomiej Boczek, Jakub Matejczyk 외 arxiv

Vision Language Models (VLMs) and Vision Language Action (VLA) models have shown promise in robotic control. Yet, they face significant challenges regarding explainability, generalization, and compute requirements. This …

PLanAR: Planning-Language-Grounded Agentic Reasoning for Robot Manipulation

2026-02-02 · Pengyuan Guo, Zhonghao Mai, Zhengtong Xu, Kaidi Zhang 외 arxiv

Recent advances in vision-language models (VLMs) have enabled increasing progress in real-world robot manipulation. However, long-horizon manipulation in unstructured environments requires VLMs to reason about changing s…

Robot Manipulation

LLM-Based Generalizable Hierarchical Task Planning and Execution for Heterogeneous Robot Teams with Event-Driven Replanning

2025-11-27 · Suraj Borate, Bhavish Rai B, Vipul Pardeshi, Madhu Vadali arxiv

This paper introduces CoMuRoS (Collaborative Multi-Robot System), a generalizable hierarchical architecture for heterogeneous robot teams that unifies centralized deliberation with decentralized execution, and supports e…

An Intelligent-Cloud Edge Multimodal Interaction System for Robots

2026-07-16 · Zihan Guo, Xiaoqi Li arxiv

Robust human-robot interaction in complex environments requires accurate gesture perception, semantic scene understanding, and reliable task planning under limited onboard computing resources. This paper presents a cloud…

Scene Understanding

Supercomputing for High-speed Avoidance and Reactive Planning in Robots

2025-09-23 · Kieran S. Lachmansingh, José R. González-Estrada, Jacob Chisholm, Ryan E. Grant 외 arxiv

This paper presents SHARP (Supercomputing for High-speed Avoidance and Reactive Planning), a proof-of-concept study demonstrating how high-performance computing (HPC) can enable millisecond-scale responsiveness in roboti…

Trajectory Planning