본문으로 건너뛰기
L8.9

RLHF: Human Preferences and Reward Models

Goal

By the end of this lesson, you can trace an RLHF-style pipeline from human preference comparisons to a reward signal and policy update, and explain why reward optimization still needs independent quality and safety evaluation.

Suppose a pretrained assistant can answer questions, but its responses are often unhelpful or poorly aligned with the behavior people want.

One practical route is to collect comparisons between candidate responses and use those comparisons to shape a training signal. This family of methods is commonly called reinforcement learning from human feedback (RLHF).

Build the pipeline from concrete artifacts​

A simplified pipeline is:

pretrained model
-> supervised fine-tuning
-> collect human preference comparisons
-> train or use a preference/reward signal
-> optimize the policy toward higher reward
-> evaluate regressions, quality, and safety

The model being optimized is the policy. Human comparisons provide evidence about which outputs are preferred under a rubric.

Preference pairs do not directly become truth​

A pair might contain:

prompt: explain the refund rule
chosen: cites the supplied 30-day policy
rejected: invents a 90-day exception

The pair says that the chosen response was preferred under the annotation criterion. It does not prove that every chosen response in the dataset is factually correct or safe.

The annotation rubric, annotator disagreement, data coverage, and shortcuts such as length all affect what the preference signal means.

What a reward model does​

A reward model or preference model can be trained to assign a scalar score that tends to rank preferred responses above rejected ones.

Conceptually:

prompt + response -> reward score

The policy can then be optimized to produce responses that receive higher predicted reward.

In historical LLM RLHF pipelines, methods such as PPO (Proximal Policy Optimization) were used for this policy-optimization stage. You do not need to implement PPO here. The important idea is that the policy is updated using a learned reward-like signal rather than a single supervised target sentence.

Why keep a reference or base constraint​

If optimization pushes only toward higher learned reward, the policy may move too far from useful pretrained behavior. RLHF systems therefore commonly use constraints or penalties that discourage excessive deviation from a reference/base policy.

The exact algorithm can differ, but the engineering question is stable: how do we improve the measured preference signal without allowing uncontrolled drift?

Reward optimization can exploit mistakes in the reward​

A reward model is imperfect. If it accidentally gives extra points to long confident answers, the policy can learn to produce longer, more confident answers even when they are wrong.

This is sometimes described as reward hacking or exploiting a proxy.

The core lesson is broader than the name:

better measured reward != guaranteed better real-world behavior

Independent evaluation is still required for factuality, task success, safety, and retention.

The bridge to DPO​

An RLHF-style route can be summarized as:

preferences -> reward model/signal -> RL policy optimization

A DPO-style route uses the preference pairs more directly:

preferences -> direct preference objective

DPO was designed to avoid an explicit reward-model-plus-RL optimization loop for this setting. That does not make DPO and RLHF universally identical or interchangeable. They are different optimization routes built from related preference evidence.

Predict

A reward model gives high scores to confident answers, including confident false answers. What should the team conclude?

Run the Browser Lab​

The deterministic Lab shows:

  • two candidate responses;
  • a preference/reward signal;
  • a policy score update toward the preferred response;
  • a second case where a positive measured reward is attached to a false response.

Complete the tiny update function. Then inspect the proxy-failure case: its positive reward moves the policy score upward even though the independent truth_ok check remains false. The goal is to see the failure happen through the same update rule, not only read a warning about it.

Loading lab…

Quick Check

1. What role does a reward model play in an RLHF-style pipeline?
2. Why use a reference/base constraint during policy optimization?
3. What is the high-level difference between the RLHF-style route and DPO-style route taught here?

0 of 3 questions answered.

Explain it back​

Trace the pipeline from a pretrained model to SFT, preference collection, reward modeling, policy optimization, and independent evaluation. Then explain one way a policy could improve reward while becoming worse on a real requirement.

Key Takeaways

  • RLHF uses human preference evidence to shape a reward-like training signal.
  • A reward model predicts preference; it does not guarantee truth or safety.
  • Policy optimization can exploit weaknesses in an imperfect reward signal.
  • Reference/base constraints help limit uncontrolled drift.
  • Independent evaluation remains necessary.
  • DPO uses a different, more direct preference-optimization route rather than an explicit reward-model-plus-RL loop.

Next Lesson

Next, study the DPO-style direct preference objective and compare its chosen-versus-rejected scoring with the RLHF-style route you just traced.

References

Lesson actions

Completion is stored locally on this device.

View progress