paper-with-me

홈 › Papers

Scaling Behavior Cloning Improves Causal Reasoning: An Open Model for Real-Time Video Game Playing

2026-01-08 · Yuguang Yue, Irakli Salia, Samuel Hunt, Chris Green, Wenzhe Shi, Jonathan J Hunt arxiv

Behavior cloning has seen a resurgence as scaling model and data sizes demonstrate strong performance. In this work, we introduce an open recipe for training a video game playing foundation model designed for inference in realtime on a consumer GPU. We release all data (8300+ hours of high quality human gameplay), training and inference code, and pretrained checkpoints under an open license. Empirically, we show that our best model achieves performance competitive with human players across a variety of 3D games. We use this recipe to investigate the scaling laws of behavior cloning, with a focus on causal reasoning. In a controlled toy setting, we first demonstrate that increasing training data and network depth leads to the model learning a more causal policy. We then validate these findings at scale, analyzing models up to 1.2 billion parameters. We observe that the causal improvements seen in the toy domain hold true as model size and training steps increase.

📄 PDF Abstract BibTeX arXiv:2601.04575

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Exploring the Limitations of Behavior Cloning for Autonomous Driving

2019-04-18 · ICCV 2019 10 · Felipe Codevilla, Eder Santana, Antonio M. López, Adrien Gaidon

Driving requires reacting to a wide variety of complex environment conditions and agent behaviors. Explicitly modeling each possible scenario is unrealistic. In contrast, imitation learning can, in theory, leverage data …

Autonomous DrivingImitation Learning

Object-Aware Regularization for Addressing Causal Confusion in Imitation Learning

2021-10-27 · NeurIPS 2021 12 · Jongjin Park, Younggyo Seo, Chang Liu, Li Zhao 외

Behavioral cloning has proven to be effective for learning sequential decision-making policies from expert demonstrations. However, behavioral cloning often suffers from the causal confusion problem where a policy relies…

Decision MakingImitation LearningSequential Decision Making

PRTS: A Primitive Reasoning and Tasking System via Contrastive Representations

2026-04-30 · Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Fangzheng Yan 외 arxiv

Vision-Language-Action (VLA) models advance robotic control via strong visual-linguistic priors. However, existing VLAs predominantly frame pretraining as supervised behavior cloning, overlooking the fundamental nature o…

Reinforcement Learning

Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models

2026-06-10 · Darpan Aswal, Thomas Palmeira Ferraz, Yongxin Zhou, Maxime Peyrard arxiv

Latent reasoning models (LRMs) replace explicit chain-of-thought with continuous thoughts. Recent work treats observable latent-state patterns, such as BFS-like frontiers and decodable arithmetic computation, as evidence…

Behavior Cloning in OpenAI using Case Based Reasoning

2020-02-23 · Chad Peters, Babak Esfandiari, Mohamad Zalat, Robert West

Learning from Observation (LfO), also known as Behavioral Cloning, is an approach for building software agents by recording the behavior of an expert (human or artificial) and using the recorded data to generate the requ…

OpenAI Gym