Reinforcement learning: actions have consequences

Some choices change what you can see and do next. A small reward now may cost a later opportunity in a game; a robot action affects whether the task can finish. RL studies learning to act from interaction feedback.

Reinforcement learning studies sequences of actions and their consequences. A model observes an environment, acts, and receives another observation and a reward. A good immediate score may lead to a poor eventual outcome. Games, robot control, and language-model post-training provide different settings for learning from such feedback.

Begin with two questions: how is an action judged, and how does that judgment change behavior? Value learning, policy optimization, and their combination in actor–critic methods offer different answers. Then examine where data comes from, whether it can be reused, and how rewards are defined.

Research tasks
  • Trade immediate gains against future opportunities, such as collecting resources before fighting in a game.
  • Learn continuous control, such as choosing the force applied at a joint.
  • Improve an existing policy using task feedback, such as whether a generated answer passes a verifier.

Write down observations, actions, rewards, and data sources before choosing an algorithm. After scores improve, inspect what the policy actually does.

How rewards shape a policy
  1. 01Exploration and reward
  2. 02State values
  3. 03Action policy
  4. 04Execution

Research milestones

  1. 2015

    Estimate the long-term value of actions

    DQN combines pixel inputs with a neural estimate of future return for each discrete action.

  2. 2017

    Improve the action-selection policy directly

    PPO updates policies with newly collected experience while discouraging overly large changes in an update.

  3. 2018

    Reuse experience while retaining exploration

    SAC combines actor and critic networks, replayed data, and an entropy objective for continuous control. These routes suit different data and action settings.

Key concepts

Policy
A rule for selecting an action from observations. It can return a fixed action or a distribution over actions.
Reward / Return
A reward is feedback at one step. Return accumulates later rewards, often discounting more distant ones.
Value
Expected future return from a state, or after taking an action, under a specified subsequent policy.
On-policy / Off-policy
On-policy updates mainly use experience from the current or nearby policy; off-policy learning can use experience from other or older policies.

States, actions, rewards and an interaction

Use Hands-on Reinforcement Learning to meet states, actions, rewards, returns, and policies. Then inspect how observations and actions flow in Gymnasium Basic Usage.

Trace a run: what did the environment provide, what did the policy choose, how did the environment change, and when did the task end? Then study how an algorithm learns from those records.

Value estimation and policy optimization

One route estimates future return from a situation or action; another directly adjusts the action-selection policy. Many algorithms combine them.

For expectations and conditional probability, revisit mathematics; for neural training, revisit DL. Choose an appropriate CleanRL implementation and connect update equations to code.

For depth, use Berkeley's deep RL course. Another Chinese explanation is Easy RL.

Research routes in reinforcement learning

Language models produce tokens, agents call tools, and humanoids control joints. These applications share algorithms but differ in data, environments and evaluation. The routes below follow research questions; methods such as model-based and offline RL can also be used in robotics and other applications.

Language-model post-training: improving responses through feedback

Once a model can generate text, how can it follow instructions or solve problems more reliably? RL can use a reward model trained on human preferences, or verifiable feedback from answers and program tests. Track where the reward comes from and whether higher reward corresponds to better responses.

  • InstructGPT: follow demonstrations, preference comparisons, reward modeling and policy optimization through the RLHF pipeline.
  • DeepSeek-R1: examine rewards for reasoning and the different training pipelines of R1-Zero and R1.
  • verl: connect response sampling, reward computation and policy updates to code. More model reports are on the LLM page.

Agentic RL: learning to complete tasks over multiple turns

An action might search, run code or operate an interface. Its result affects the next decision. Training must account for an entire interaction: which actions contributed to success, how errors accumulated, and whether tool use actually helped.

  • Search-R1: integrates search into reasoning and uses RL to learn when to retrieve and how to use results; see the project code.
  • RAGEN: studies training stability, environment feedback and policy learning in multi-turn interactions.

Trace a full trajectory: which text came from the model, which observations came from the environment, and when rewards arrived. Tool use and task execution are introduced on the Agent page.

Humanoid control: from simulation to real robots

Balancing, walking, motion tracking and whole-body coordination require continuous control. Policies are often trained in simulation and transferred to hardware, where contact, latency and dynamics differ. Examine observations, action representations, rewards and how simulation is aligned with reality.

  • Humanoid-Gym: introduces simulation training and zero-shot sim-to-real transfer for humanoid locomotion.
  • ASAP: aligns simulated and real dynamics to learn agile whole-body skills.
  • Unitree RL Gym: follow robot configurations, training environments and deployment code. For manipulation and VLA models, continue to embodied AI.

Model-based RL: predicting the consequences of actions

Learn how an environment changes under actions, then use that model to train a policy or plan. Predictions can live in latent space without producing clear images. Distinguish learning a policy from imagined trajectories from searching over candidate actions during execution.

  • DreamerV3: learns a world model and trains behavior in imagined trajectories; follow the connection between model learning and policy learning.
  • TD-MPC2: learns a latent model for control and selects actions through model predictive control; compare its decision process with Dreamer's.

For broader work on prediction, video simulation and interactive environments, see world models.

Offline RL: learning from existing interaction data

When real-world trials are expensive or only previously collected data is available, a policy must learn from a fixed dataset. Behavior cloning imitates recorded actions. Offline RL also uses rewards and long-term returns, while addressing unreliable value estimates when a policy moves beyond the data distribution.

  • CQL: uses conservative value estimation to reduce overoptimism about actions poorly covered by the data.
  • IQL: combines implicit value learning with advantage-weighted policy extraction, reducing reliance on values of out-of-distribution actions during training.

Check which policies collected the data, which behaviors it covers, and how improvements beyond the demonstrations are evaluated.

Multi-agent RL: coordinating multiple decision-makers

When several agents act together, each may see only part of the environment while the others' policies also change. Cooperation, competition, communication and credit assignment become central questions. A shared team reward makes individual contributions particularly hard to identify.

  • MAPPO: examines PPO in cooperative multi-agent tasks, including centralized value estimation and decentralized execution.
  • Official implementation: compare local observations, global information and individual actions to understand what information is available during training and execution.

LLM multi-agent systems also involve coordination; whether they use RL depends on whether their policies are trained through interaction rewards.

Change the reward and observe the policy

Keep configuration fixed in a small environment and repeat runs. Compare training curves and final behavior. Then change a reward setting and look for unexpected behavior. For unusual results, inspect actual trajectories first.

RL textbooks, implementations and projects

Chinese textbooks, small environments, implementations, and advanced courses are in the RL catalog. For robotics and control projects, see labs.

Representative papers

Start with interaction and updates in DQN, then compare how PPO and SAC use data. Explaining what each learns and which data it needs is a useful first reading.

2015 · Nature

Human-level control through deep reinforcement learning

Human-level control through deep reinforcement learning · Mnih et al. · Fig. 1
Figure excerpt from the paper · Mnih et al. · Fig. 1 · Original paper

The problem

How can pixels guide discrete actions?

The key idea

Predict action values; replay experience and use a delayed target network to stabilize learning.

Why this paper

A clear entry to observations, values, and deep RL training targets.

Where to start

Trace the interaction loop, Q updates, and Atari evaluation.

A question to keep asking

The original method enumerates discrete actions; continuous control needs other machinery.

2017 · arXiv

Proximal Policy Optimization Algorithms

Proximal Policy Optimization Algorithms · Schulman et al. (2017), Fig. 1
Figure excerpt from the paper · Schulman et al. (2017), Fig. 1 · Original paper

The problem

How can simple policy updates avoid overly large changes?

The key idea

The clipped objective removes incentives for excessive changes in action-probability ratios.

Why this paper

A practical entry to policy gradients and many RL implementations.

Where to start

Inspect the clipped objective and the sampling/minibatch-update loop.

A question to keep asking

Clipping does not guarantee monotonic improvement; collection and implementation choices still matter.

2018 · ICML

Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor · Haarnoja et al. (2018), Fig. 1
Figure excerpt from the paper · Haarnoja et al. (2018), Fig. 1 · Original paper

The problem

How can experience be reused without committing too early to one action?

The key idea

Optimize reward and policy entropy, using a critic and replayed experience.

Why this paper

Compare with PPO to understand data reuse and exploration objectives.

Where to start

Connect the maximum-entropy objective, actor, critic, and replay buffer.

A question to keep asking

This paper focuses on continuous control; reward and entropy scales affect behavior.

Getting started

Stuck on a step? Bring your attempt to the AMA ↗