Train and evaluate policies that select actions from observations and rewards. Start in a simulator, compare against a simple policy, and report behavior across
Last reviewed: 2026-10-03
Reinforcement learning studies policies that choose actions from observations and receive rewards from an environment. Gymnasium defines a practical interface for observations, actions, rewards, resets, and episode outcomes. It distinguishes task termination from external truncation such as a time limit.
Start with a small simulated environment and a random or rule-based policy. Evaluate repeated runs and behavior outside the training scenarios. A policy can improve its reward while doing something undesirable if the reward or environment omits an important requirement.
No. Simulator performance is evidence about that environment. Physical deployment requires a separate engineering and safety-validation process.