본문으로 건너뛰기
L8.7

Reinforcement Learning: States, Actions, Rewards, and Policies

Goal

By the end of this lesson, you can explain the reinforcement-learning interaction loop using agent, environment, state or observation, action, reward, policy, return, and exploration versus exploitation.

Imagine a game in which you repeatedly choose between two buttons. After each choice, the game gives you a number. One button often gives a small reward. The other sometimes gives a larger reward but sometimes gives nothing.

You are not given the best action as a label for every round. You act, observe what happens, receive a reward, and use experience to improve future choices.

That repeated feedback loop is the core of reinforcement learning (RL).

The basic interaction loop​

The pieces are:

  • agent — the learner or decision-maker;
  • environment — the world the agent interacts with;
  • state/observation — information available before a decision;
  • action — a choice the agent makes;
  • reward — a feedback number after the action;
  • policy — the rule or model that chooses actions.

The loop repeats:

observe state
-> choose action from policy
-> environment changes
-> receive reward and next observation
-> repeat

In a simple bandit there may be no changing state; the important uncertainty is which action tends to produce better reward. In other tasks, the state changes after every action.

Reward is feedback, not a complete description of the world​

A reward tells the agent what the training process values numerically. It does not automatically capture everything we care about.

If a delivery robot receives +1 only for moving quickly, it may learn behavior that is fast but unsafe unless the reward or constraints represent safety too.

That distinction will matter again in RLHF: optimizing a measured reward does not prove truth, safety, or general quality.

Longer-term consequences need return​

Some actions produce a small immediate reward but a better future. RL therefore often reasons about return: accumulated reward over several steps, sometimes giving less weight to rewards far in the future.

You do not need the Bellman equation here. The intuitive question is enough: does this action only look good now, or does it lead to better outcomes over the rest of the interaction?

Exploration versus exploitation​

If the agent always chooses the action that currently looks best, it may never discover a better option. If it explores constantly, it may ignore useful knowledge it already has.

  • exploitation uses the best-known action;
  • exploration tries alternatives to gather information.

A practical policy often balances both.

Two major families you will hear about​

Two important names are:

  • Q-learning/value-based methods, which learn estimates of how good actions or states are;
  • policy-gradient methods, which directly adjust a parameterized policy toward actions associated with higher return.

These names locate later methods on the map. This lesson does not require implementing a full RL algorithm.

Predict

An agent always chooses the action with the highest reward estimate and never tries alternatives. Which side of the trade-off is it emphasizing?

Run the Browser Lab​

The Lab is a deterministic two-action bandit.

  1. Run an action sequence that chooses only the safe action and inspect its per-round rewards and total return.
  2. Try an action sequence that explores the risky action on selected rounds.
  3. Compare the total return and the information learned from the two sequences. A sequence is a recorded set of choices; a policy is the rule that would generate choices from observations.
  4. Change one action at a time and predict the reward before running.
  5. Explain why a single lucky reward should not be confused with a guaranteed best policy.

Loading lab…

Quick Check

1. What is a policy?
2. Why can return differ from immediate reward?
3. Why explore?

0 of 3 questions answered.

Explain it back​

Describe an agent interacting with an environment for three steps. Name the observation, action, reward, and policy at each step, then explain one reason an immediate reward could be misleading.

Key Takeaways

  • RL learns from repeated interaction rather than target labels for every decision.
  • A policy chooses actions; the environment returns observations and rewards.
  • Return captures longer-term consequences beyond the next reward.
  • Exploration gathers information; exploitation uses what is currently believed to work.
  • Reward is a training signal, not proof that every desired property is satisfied.

Next Lesson

Next, return to language-model preference pairs. Those comparisons will become the human feedback used by RLHF-style and DPO-style alignment methods.

References

Lesson actions

Completion is stored locally on this device.

View progress