Record summary
A quick snapshot of what this page covers.
Risk profile
How this risk is described and categorized.
"Limitations of Reward Modeling. Training reward models using comparison feedback can pose significantchallenges in accurately capturing human values. For example, these models may unconsciously learn suboptimal or incomplete objectives, resulting in reward hacking (Zhuang and Hadfield-Menell, 2020; Skalse et al.,2022). Meanwhile, using a single reward model may struggle to capture and specify the values of a diversehuman society (Casper et al., 2023b)."
Suggested mitigations
Defenses that may help with related attacks.
Source
Research source for this risk, when available.
Included resource
AI Alignment: A Comprehensive Survey
Original source
MIT AI Risk Repository
Open the public repository used for AI risk records and taxonomy fields.
