paper-with-me

홈 › Papers

vla.cpp: A Unified Inference Runtime for Vision-Language-Action Models

2026-06-06 · Khanh D. Nguyen, Hung T. Ho, Chinh T. Nguyen, Thanh Q. Duong, Linh D. Le, Duy M. H. Nguyen, Vien A. Ngo, An T. Le arxiv

Vision-Language-Action (VLA) policies are typically shipped as Python/PyTorch stacks that assume a workstation-class GPU, a mismatch for the hardware on which robots actually run. We present vla.cpp, a portable C++ inference runtime built on llama.cpp. To our knowledge, it is the first ggml-class engine to natively serve the flow-matching and diffusion VLA inference pattern, in which a cached vision-language prefix is consumed by a cross-attending action expert integrated over several solver steps. A single runtime serves seven architectures spanning five backbone and four action-head families behind one request/response protocol, with each model packaged as a self-contained bundle. On LIBERO-Object, the engine matches a state-of-the-art checkpoint to within one episode out of 200, and runs BitVLA at 100% success in 1.3 GiB of memory. The same bundle runs unchanged across three hardware tiers, from a consumer GPU down to an 8 GB embedded module. A cross-hardware roofline analysis shows that batch-1 VLA inference is compute-bound, so utilization rather than bandwidth is the deployment lever; an IMMA ladder GEMM derived from this analysis cuts BitVLA per-step latency by 4.5x. We then frame an on-robot stress test on an ALOHA arm that isolates the latency constraint under which a learned VLA must replan against a moving target on the hardware it was trained for. Code, demo videos, and the reproducible benchmark scaffold are available at https://fai-modelopt-tech.github.io/vla-cpp.github.io/.

📄 PDF Abstract BibTeX arXiv:2606.08094

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms

2026-04-26 · Qi Li, Bo Yin, Weiqi Huang, Ruhao Liu 외 arxiv

Vision-Language-Action (VLA) models are emerging as a unified substrate for embodied intelligence. This shift raises a new class of safety challenges, stemming from the embodied nature of VLA systems, including irreversi…

LiteVLA-Edge: Quantized On-Device Multimodal Control for Embedded Robotics

2026-03-03 · Justin Williams, Kishor Datta Gupta, Roy George, Mrinmoy Sarkar arxiv

Vision-Language-Action (VLA) models provide a unified framework for perception, language conditioning, and action generation, but many existing systems remain difficult to deploy in embedded robotic settings because of t…

BaseRT: Best-in-Class LLM Inference on Apple Silicon via Native Metal

2026-07-01 · Prabod Rathnayaka, Fabian Waschkowski, Lukas Wesemann arxiv

We present BaseRT, a native Metal inference runtime for large language models (LLMs) on Apple Silicon, and report the highest inference throughput on this hardware to date. Existing runtimes, including llama.cpp and MLX-…

TRINE: A Token-Aware, Runtime-Adaptive FPGA Inference Engine for Multimodal AI

2026-03-24 · Hyunwoo Oh, Hanning Chen, Sanggeon Yun, Yang Ni 외 arxiv

Multimodal stacks that mix ViTs, CNNs, GNNs, and transformer NLP strain embedded platforms because their compute/memory patterns diverge and hard real-time targets leave little slack. TRINE is a single-bitstream FPGA acc…

Quantization with Unified Adaptive Distillation to enable multi-LoRA based one-for-all Generative Vision Models on edge

2026-03-31 · Sowmya Vajrala, Aakash Parmar, Prasanna R, Sravanth Kodavanti 외 arxiv

Generative Artificial Intelligence (GenAI) features such as image editing, object removal, and prompt-guided image transformation are increasingly integrated into mobile applications. However, deploying Large Vision Mode…

Image Editing