paper-with-me

홈 › Papers

Dexbotic: Open-Source Vision-Language-Action Toolbox

2025-10-27 · Bin Xie, Erjin Zhou, Fan Jia, Hao Shi, Haoqiang Fan, Haowei Zhang, Hebei Li, Jianjian Sun, Jie Bin, Junwen Huang, Kai Liu, Kaixin Liu, Kefan Gu, Lin Sun, Meng Zhang, Peilong Han, Ruitao Hao, Ruitao Zhang, Saike Huang, Songhan Xie, Tiancai Wang, Tianle Liu, Wenbin Tang, Wenqi Zhu, Yang Chen, Yingfei Liu, Yizhuang Zhou, Yu Liu, Yucheng Zhao, Yunchao Ma, Yunfei Wei, Yuxiang Chen, Ze Chen, Zeming Li, Zhao Wu, Ziheng Zhang, Ziming Liu, Ziwei Yan, Ziyu Zhang arxiv

In this paper, we present Dexbotic, an open-source Vision-Language-Action (VLA) model toolbox based on PyTorch. It aims to provide a one-stop VLA research service for professionals in the field of embodied intelligence. It offers a codebase that supports multiple mainstream VLA policies simultaneously, allowing users to reproduce various VLA methods with just a single environment setup. The toolbox is experiment-centric, where the users can quickly develop new VLA experiments by simply modifying the Exp script. Moreover, we provide much stronger pretrained models to achieve great performance improvements for state-of-the-art VLA policies. Dexbotic will continuously update to include more of the latest pre-trained foundation models and cutting-edge VLA models in the industry.

📄 PDF Abstract BibTeX arXiv:2510.23511

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models

2025-06-10 · Pranav Guruprasad, Yangyue Wang, Sudipta Chowdhury, Jaewoo Song 외

Recent innovations in multimodal action models represent a promising direction for developing general-purpose agentic systems, combining visual understanding, language comprehension, and action generation. We introduce M…

Action GenerationImage CaptioningQuestion AnsweringVision-Language-Action+1

OpenDriveVLA: Towards End-to-end Autonomous Driving with Large Vision Language Action Model

2025-03-30 · Xingcheng Zhou, Xuyuan Han, Feng Yang, Yunpu Ma 외

We present OpenDriveVLA, a Vision-Language Action (VLA) model designed for end-to-end autonomous driving. OpenDriveVLA builds upon open-source pre-trained large Vision-Language Models (VLMs) to generate reliable driving …

Autonomous DrivingDecision MakingMotion PlanningQuestion Answering+4

Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models

2025-05-29 · Haohan Chi, Huan-ang Gao, Ziming Liu, Jianing Liu 외

Vision-Language-Action (VLA) models for autonomous driving show promise but falter in unstructured corner case scenarios, largely due to a scarcity of targeted benchmarks. To address this, we introduce Impromptu VLA. Our…

Autonomous DrivingDiagnosticQuestion AnsweringTrajectory Prediction+1

Typhoon OCR: Open Vision-Language Model For Thai Document Extraction

2026-01-21 · Surapon Nonesung, Natapong Nitarach, Teetouch Jaknamon, Pittawat Taveekitworachai 외 arxiv

Document extraction is a core component of digital workflows, yet existing vision-language models (VLMs) predominantly favor high-resource languages. Thai presents additional challenges due to script complexity from non-…

ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai

2025-11-06 · Surapon Nonesung, Teetouch Jaknamon, Sirinya Chaiophat, Natapong Nitarach 외 arxiv

We present ThaiOCRBench, the first comprehensive benchmark for evaluating vision-language models (VLMs) on Thai text-rich visual understanding tasks. Despite recent progress in multimodal modeling, existing benchmarks pr…