--- aliases: en: - reward-hacking ja: - リワードハッキング concept_id: reward_hacking preferred: en: reward hacking ja: 報酬ハッキング relations: relations: prerequisite: - reward status: active --- # reward hacking A family of failures in which a system obtains or increases reward through an unintended route rather than by achieving the designer's intended outcome. This corpus uses reward hacking broadly enough to cover both specification gaming, where the specified reward is exploited, and reward tampering, where the mechanism that computes, records, communicates, or delivers reward is manipulated; some literature instead uses 'reward hacking' more narrowly as a near-synonym for specification gaming. ## prerequisite - [reward](../terminology/reward.txt) ## Documents: about - [チャット形式の安全性訓練で望ましく振る舞うことは、エージェント的な状況での安全性を保証しない](../documents/macdiarmid_et_al.chat-alignment-does-not-establish-agentic-alignment.txt) ## Documents: mentions - [AIコントロールは時間を稼げても、超知能に対する長期的な安全計画としては頼れない](../documents/greenblatt.control-can-buy-time-but-is-not-a-superintelligence-safety-plan.txt) [Corpus index](../index.txt)