Preference Optimization Concepts
Goal
Explain preference optimization as increasing the model's relative preference for chosen responses over rejected responses, compute a toy preference margin, and distinguish preference fit from general model quality.
Suppose two answers are shown for the same question. A reviewer marks answer A as preferable to answer B because A stays inside the evidence. Preference training does not need to treat A as the one perfect sentence for all time. The useful signal is the comparison: under this same context, the model should score A above B.
The numerical margin later in this Lesson is just a ruler for that ordering. If the chosen answer receives a higher model score than the rejected answer, the margin is positive. If the rejected answer is scored higher, the margin is negative. Once that direction is clear, log-probabilities give us a precise way to calculate it.
Preference optimization uses pairwise evidence:
prompt x
chosen response y+
rejected response y-
The core training goal is to make the model favor y+ more strongly than y- under the same context.
Start with log-probability scores
Suppose the current model assigns:
log p(chosen | x) = -4.0
log p(rejected | x) = -3.0
The rejected response currently has the higher score because -3.0 > -4.0.
A simple margin is:
margin = logp_chosen - logp_rejected
= -4.0 - (-3.0)
= -1.0
Negative margin means the model currently prefers the rejected candidate. After adaptation:
chosen = -2.5
rejected = -3.5
margin = +1.0
Now the relative ordering matches the pair label.
DPO-style objectives use relative changes
Direct Preference Optimization uses a more specific objective involving the policy model, a reference policy, and a logistic form. You do not need to memorize the full equation before understanding the design idea:
- compare chosen versus rejected;
- compare that preference relative to a reference model;
- optimize a smooth objective that rewards the desired relative preference.
The reference term helps anchor the adapted policy to the starting behavior.
Sequence scores need careful interpretation
A response contains multiple tokens. Sequence log probability is often a sum over token log probabilities:
log p(y|x) = Σ log p(token_t | x, previous tokens)
Longer sequences accumulate more negative terms. This means raw sequence score can interact with response length. Real implementations may use specific normalization or objective conventions. Check the library/paper rather than inventing an unrecorded rule.
Preference fit is not the same as truth
A model can learn to prefer the chosen responses and still:
- hallucinate on unseen questions;
- regress on unrelated capabilities;
- overfit annotator style;
- exploit superficial dataset shortcuts.
Preference loss measures fit to the pairwise objective. It does not replace factual, safety, task, or regression evaluation.
One pair can move for the wrong reason
Suppose chosen responses are always shorter. The model may improve preference accuracy partly by learning a general short-response bias. That can raise the preference metric while harming tasks where detailed answers are required. This is why preference-data auditing and downstream evaluation belong in the same workflow.
Preference fit is relative, not a universal quality score
A preference objective compares alternatives under a context and criterion. Raising the chosen response relative to the rejected response means the model is becoming more consistent with those pairs. It does not prove that either answer is factually correct, safe, or useful outside the represented preference distribution.
Keep an independent task-quality or retention suite beside preference metrics. Otherwise a model can become better at the measured pairwise objective while drifting on unrelated capabilities. The reference policy used in DPO-style reasoning also matters because it anchors how far the adapted policy moves relative to a baseline rather than treating raw chosen likelihood as the only quantity of interest.
Predict
Complete the Browser Lab objective
The Browser Lab contains pair scores before and after a toy adaptation.
- Click Run once. Every margin is
0.0, so the checks fail. - Complete
preference_marginso it returns how much higher the chosen response scores than the rejected one. - Click Run again. Pair
ashould showmargin= 3.0 chosen_preferred= True, and pairbmargin= -1.0 chosen_preferred= False. - Add a pair where both scores are good, but the rejected response is scored even higher:
{"id": "c", "chosen_logp": -1.5, "rejected_logp": -0.5, "quality_ok": True},
- Before running, predict the margin.
- Click Run. Pair
cshowsmargin= -1.0andchosen_preferred= False, even thoughquality_ok= True. The preference objective and an independent task-quality check can disagree, so you need both.
Loading lab…
Quick Check
Explain it back
Use two log-probability scores to calculate a preference margin. Then give one example where the margin improves but the system still needs another evaluation metric.
Key Takeaways
- Preference optimization changes relative chosen/rejected scoring.
- Sequence log probabilities aggregate token-level probabilities.
- DPO-style objectives compare policy preference relative to a reference.
- Preference fit is not proof of factual correctness or broad capability.
- Dataset shortcuts can improve the objective for the wrong reason.
Next Lesson
Next, compare the base and adapted model on the same fixed cases so improvements and regressions are visible together.
References
- Rafailov et al., Direct Preference Optimization.
Completion is stored locally on this device.