paper-with-me

Papers

CombatVLA: An Efficient Vision-Language-Action Model for Combat Tasks in 3D Action Role-Playing Games

2025-03-12 · Peng Chen, Pi Bu, Yingyao Wang, Xinyi Wang, ZiMing Wang, Jie Guo, Yingxiu Zhao, Qi Zhu, Jun Song, Siran Yang, Jiamang Wang, Bo Zheng

Recent advances in Vision-Language-Action models (VLAs) have expanded the capabilities of embodied intelligence. However, significant challenges remain in real-time decision-making in complex 3D environments, which demand second-level responses, high-resolution perception, and tactical reasoning under dynamic conditions. To advance the field, we introduce CombatVLA, an efficient VLA model optimized for combat tasks in 3D action role-playing games(ARPGs). Specifically, our CombatVLA is a 3B model trained on video-action pairs collected by an action tracker, where the data is formatted as action-of-thought (AoT) sequences. Thereafter, CombatVLA seamlessly integrates into an action execution framework, allowing efficient inference through our truncated AoT strategy. Experimental results demonstrate that CombatVLA not only outperforms all existing models on the combat understanding benchmark but also achieves a 50-fold acceleration in game combat. Moreover, it has a higher task success rate than human players. We will open-source all resources, including the action tracker, dataset, benchmark, model weights, training code, and the implementation of the framework at https://combatvla.github.io/.

📄 PDF Abstract BibTeX arXiv:2503.09527

Code (1)

ChenVoid/CombatVLA 공식 구현

Tasks

Decision MakingVision-Language-Action

Similar Papers 제목 키워드 기반

Cultivating Game Sense for Yourself: Making VLMs Gaming Experts

2025-03-27 · Wenxuan Lu, Jiangyang He, Zhanqiu Zhang, Yiwen Guo 외

Developing agents capable of fluid gameplay in first/third-person games without API access remains a critical challenge in Artificial General Intelligence (AGI). Recent efforts leverage Vision Language Models (VLMs) as d…

Can VLMs Play Action Role-Playing Games? Take Black Myth Wukong as a Study Case

2024-09-19 · Peng Chen, Pi Bu, Jun Song, Yuan Gao 외

Recently, large language model (LLM)-based agents have made significant advances across various fields. One of the most popular research areas involves applying these agents to video games. Traditionally, these methods h…

Large Language Model

A Hierarchical Deep Reinforcement Learning Framework for 6-DOF UCAV Air-to-Air Combat

2022-12-05 · Jiajun Chai, Wenzhang Chen, Yuanheng Zhu, Zong-xin Yao 외

Unmanned combat air vehicle (UCAV) combat is a challenging scenario with continuous action space. In this paper, we propose a general hierarchical framework to resolve the within-vision-range (WVR) air-to-air combat prob…

Deep Reinforcement LearningReinforcement Learning (RL)

Combatting Human Trafficking in the Cyberspace: A Natural Language Processing-Based Methodology to Analyze the Language in Online Advertisements

2023-11-22 · Alejandro Rodriguez Perez, Pablo Rivas

This project tackles the pressing issue of human trafficking in online C2C marketplaces through advanced Natural Language Processing (NLP) techniques. We introduce a novel methodology for generating pseudo-labeled datase…

Action DetectionActivity Detection

Research on Autonomous Maneuvering Decision of UCAV based on Approximate Dynamic Programming

2019-08-27 · Zhencai Hu, Peng Gao, Fei Wang

Unmanned aircraft systems can perform some more dangerous and difficult missions than manned aircraft systems. In some highly complicated and changeable tasks, such as air combat, the maneuvering decision mechanism is re…

Decision Makingfeature selection