Reinforcement Learning: States, Actions, Rewards, and Policies
Goal
By the end of this lesson, you can explain the reinforcement-learning interaction loop using agent, environment, state or observation, action, reward, policy, return, and exploration versus exploitation.
Imagine a game in which you repeatedly choose between two buttons. After each choice, the game gives you a number. One button often gives a small reward. The other sometimes gives a larger reward but sometimes gives nothing.
You are not given the best action as a label for every round. You act, observe what happens, receive a reward, and use experience to improve future choices.
That repeated feedback loop is the core of reinforcement learning (RL).
The basic interaction loop
The pieces are:
- agent — the learner or decision-maker;
- environment — the world the agent interacts with;
- state/observation — information available before a decision;
- action — a choice the agent makes;
- reward — a feedback number after the action;
- policy — the rule or model that chooses actions.
The loop repeats:
observe state
-> choose action from policy
-> environment changes
-> receive reward and next observation
-> repeat
In a simple bandit there may be no changing state; the important uncertainty is which action tends to produce better reward. In other tasks, the state changes after every action.
Reward is feedback, not a complete description of the world
A reward tells the agent what the training process values numerically. It does not automatically capture everything we care about.
If a delivery robot receives +1 only for moving quickly, it may learn behavior that is fast but unsafe unless the reward or constraints represent safety too.
That distinction will matter again in RLHF: optimizing a measured reward does not prove truth, safety, or general quality.
Longer-term consequences need return
Some actions produce a small immediate reward but a better future. RL therefore often reasons about return: accumulated reward over several steps, sometimes giving less weight to rewards far in the future.
You do not need the Bellman equation here. The intuitive question is enough: does this action only look good now, or does it lead to better outcomes over the rest of the interaction?
Exploration versus exploitation
If the agent always chooses the action that currently looks best, it may never discover a better option. If it explores constantly, it may ignore useful knowledge it already has.
- exploitation uses the best-known action;
- exploration tries alternatives to gather information.
A practical policy often balances both.
Two major families you will hear about
Two important names are:
- Q-learning/value-based methods, which learn estimates of how good actions or states are;
- policy-gradient methods, which directly adjust a parameterized policy toward actions associated with higher return.
These names locate later methods on the map. This lesson does not require implementing a full RL algorithm.
Predict
Run the Browser Lab
The Lab is a deterministic two-action bandit.
- Run an action sequence that chooses only the safe action and inspect its per-round rewards and total return.
- Try an action sequence that explores the risky action on selected rounds.
- Compare the total return and the information learned from the two sequences. A sequence is a recorded set of choices; a policy is the rule that would generate choices from observations.
- Change one action at a time and predict the reward before running.
- Explain why a single lucky reward should not be confused with a guaranteed best policy.
Loading lab…
Quick Check
Explain it back
Describe an agent interacting with an environment for three steps. Name the observation, action, reward, and policy at each step, then explain one reason an immediate reward could be misleading.
Key Takeaways
- RL learns from repeated interaction rather than target labels for every decision.
- A policy chooses actions; the environment returns observations and rewards.
- Return captures longer-term consequences beyond the next reward.
- Exploration gathers information; exploitation uses what is currently believed to work.
- Reward is a training signal, not proof that every desired property is satisfied.
Next Lesson
Next, return to language-model preference pairs. Those comparisons will become the human feedback used by RLHF-style and DPO-style alignment methods.
References
- Sutton and Barto, Reinforcement Learning: An Introduction.
Completion is stored locally on this device.