Reinforcement learning trains agents through trial and reward: act in an environment, receive feedback, adjust toward strategies that maximize long-term reward.
Training agents through trial and reward: act in an environment, receive feedback, and adjust toward strategies that maximize long-term reward. It mastered games and robotics control, and, as RLHF and successors, became the technique that aligns language models.
The frontier applies RL to reasoning and agency: training models on verified task outcomes, code that passes tests, math that checks out, is a core engine behind reasoning models and capable agents.