paper-with-me

홈 › Papers

Solving the Rubik's Cube Without Human Knowledge

2018-05-18 · Stephen McAleer, Forest Agostinelli, Alexander Shmakov, Pierre Baldi

A generally intelligent agent must be able to teach itself how to solve problems in complex domains with minimal human supervision. Recently, deep reinforcement learning algorithms combined with self-play have achieved superhuman proficiency in Go, Chess, and Shogi without human data or domain knowledge. In these environments, a reward is always received at the end of the game, however, for many combinatorial optimization environments, rewards are sparse and episodes are not guaranteed to terminate. We introduce Autodidactic Iteration: a novel reinforcement learning algorithm that is able to teach itself how to solve the Rubik's Cube with no human assistance. Our algorithm is able to solve 100% of randomly scrambled cubes while achieving a median solve length of 30 moves -- less than or equal to solvers that employ human domain knowledge.

📄 PDF Abstract BibTeX arXiv:1805.07470

Code (9)

AashrayAnand/rubiks-cube-reinforcement-learning
Dalkio/RL_rubiks
JasperBusschers/multi-objective-Rubik-s-cube pytorch
KeatonMueller/rl-cube pytorch
kaletap/deep-cube-rl pytorch
mkovalski/cs4995_cube tf
nascarsayan/dl-rubiks-autodidactic-solver tf
nathangrinsztajn/rubiks_cube
robbiejones96/RubiksSolver tf

Tasks

Combinatorial OptimizationDeep Reinforcement Learningreinforcement-learningReinforcement LearningReinforcement Learning (RL)Rubik's Cube

Similar Papers 제목 키워드 기반

Solving Rubik's Cube Without Tricky Sampling

2024-11-29 · Yicheng Lin, Siyu Liang

The Rubiks Cube, with its vast state space and sparse reward structure, presents a significant challenge for reinforcement learning (RL) due to the difficulty of reaching rewarded states. Previous research addressed this…

Policy Gradient MethodsReinforcement Learning (RL)Rubik's Cube

Solving the Rubik's Cube with Approximate Policy Iteration

2019-05-01 · ICLR 2019 5 · Stephen McAleer, Forest Agostinelli, Alexander Shmakov, Pierre Baldi

Recently, Approximate Policy Iteration (API) algorithms have achieved super-human proficiency in two-player zero-sum games such as Go, Chess, and Shogi without human data. These API algorithms iterate between two policie…

Rubik's Cube

CubeRobot: Grounding Language in Rubik's Cube Manipulation via Vision-Language Model

2025-03-25 · Feiyang Wang, Xiaomin Yu, Wangyu Wu

Proving Rubik's Cube theorems at the high level represents a notable milestone in human-level spatial imagination and logic thinking and reasoning. Traditional Rubik's Cube robots, relying on complex vision systems and f…

Decision MakingLanguage ModelingLanguage ModellingRubik's Cube

Sub-Optimal Multi-Phase Path Planning: A Method for Solving Rubik's Revenge

2016-01-20 · Jared Weed

Rubik's Revenge, a 4x4x4 variant of the Rubik's puzzles, remains to date as an unsolved puzzle. That is to say, we do not have a method or successful categorization to optimally solve every one of its approximately $7.40…

Rubik's CubeTime SeriesTime Series Analysis

Universality in Collective Intelligence on the Rubik's Cube

2025-11-23 · David Krakauer, Gülce Kardeş, Joshua Grochow arxiv

Progress in understanding expert performance is limited by the scarcity of quantitative data on long-term knowledge acquisition and deployment. Here we use the Rubik's Cube as a cognitive model system existing at the int…