Reward
A reward is a feedback signal used to guide learning or assess outcomes against an objective. In language-model training, a learned reward model can assign scores based on preference examples.
[Hugging Face]In practice · hypothetical example
A scorer gives one candidate response a higher value after learning from response preferences.
[Hugging Face]A little deeper
A preference-trained scorer learns to rank outputs from chosen and rejected examples. Its score is an estimate under that training setup, not an objective certificate of truth. [Hugging Face]
A common mix-up
A high reward proves an answer is true.
A reward reflects the scoring objective and can miss important qualities. [Hugging Face]
What does a learned reward score represent?
Sources & editorial notes
Evidence: supported. Primary-source support for this scoped entry; publication approved by the project owner.
- Reward Modeling ↗ (opens in new tab)Hugging Face · Publication date unknown
Relevant section: Overview; Dataset format
Last editorial review: 2026-09-13 by project-owner.
First observed in this corpus: Unknown.
Revision history
Revision 2 · Created 2026-09-13 · Updated 2026-09-13
Project owner approved the current content for publication. Existing evidence scope and limitations remain applicable.