Continuous Control
73개 벤치마크 · 논문 1,431편 · 이 태스크의 논문 보기 →
Benchmarks
PyBullet Ant
PyBullet HalfCheetah
PyBullet Hopper
PyBullet Walker2D
Lunar Lander (OpenAI Gym)
cartpole.balance_sparse
cartpole.swingup
cheetah.run
finger.turn_hard
walker.stand
walker.walk
2D Walker
Acrobot
Ant
Ant + Gathering
Ant + Maze
Cart Pole (OpenAI Gym)
Cart-Pole Balancing
Double Inverted Pendulum
Full Humanoid
Half-Cheetah
Hopper
Inverted Pendulum
Mountain Car
Simple Humanoid
Swimmer
Swimmer + Gathering
Swimmer + Maze
acrobot.swingup
ball_in_cup.catch
cartpole.balance
cartpole.swingup_sparse
finger.spin
finger.turn_easy
fish.swim
hopper.hop
hopper.stand
humanoid.run
manipulator.insert_ball
manipulator.insert_peg
pendulum.swingup
quadruped.run
quadruped.walk
reacher.easy
reacher.hard
walker.run
Most implemented
Proximal Policy Optimization Algorithms
Continuous control with deep reinforcement learning
Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
Addressing Function Approximation Error in Actor-Critic Methods
Multi-Goal Reinforcement Learning: Challenging Robotics Environments and Request for Research
Simple random search provides a competitive approach to reinforcement learning
Papers
Controllable Affective Generation via Latent Vector Steering
Large Language Models (LLMs) often produce emotionally flattened responses after alignment, limiting their effectiveness in affect-sensitive applications. In this paper, we propose EmoVec, a lightweight framework for con…
Continuous ControlGraph-Operator World Models for Morphology-Parameter Generalization in Continuous Control
World models for continuous control are commonly trained for a fixed physical system and can degrade when known morphology parameters such as link lengths, masses, damping, and actuation change. Existing approaches often…
Continuous ControlOrthogonal JEPA: Factorized Predictive States for Latent World Models
World models construct latent states that support prediction, planning, and reasoning about an underlying system. Joint-embedding predictive architectures (JEPAs) offer a direct way to learn such states by predicting tar…
Continuous ControlTo Go Far, Go Together: Diverse Preferences Induce a Curriculum for Reward Optimization
Learning a reward model from human feedback and optimizing a policy against it is one approach to aligning AI systems with individual users. From a fairness perspective, existing work improves such alignment by developin…
Continuous ControlAn Omitted Mode Is a Rare Rule: The Sampling-Verification Danger Law in Continuous Code World Models
In the Code World Model paradigm an LLM synthesizes an executable world model that a classical planner searches, and the model is accepted when it reproduces sampled transitions. We ask what that acceptance certifies in …
Continuous ControlObservation-Grounded Self-Predictive Reinforcement Learning for Visual Continuous Control
Sample-efficient policy learning from pixels is a long-standing challenge in reinforcement learning (RL). Recent dynamics-based representation learning methods have significantly improved the sample efficiency of model-f…
Representation LearningReinforcement LearningContinuous Control