--- aliases: en: - outer alignment problem - reward misspecification problem concept_id: outer_alignment preferred: en: outer alignment ja: アウターアラインメント relations: relations: broader: - ai_alignment prerequisite: - ai_alignment - alignment_target - training_objective related: - reinforcement_learning_from_human_feedback status: active --- # outer alignment The problem or property of specifying a training objective that adequately captures the designers' intended goal or broader alignment target, assuming the learned system then robustly pursues that specification. The training objective may itself be implemented through a loss, reward-derived return, preference feedback, or another criterion. The inner/outer distinction is useful but its exact boundary varies across authors and settings. ## broader - [AI alignment](../terminology/ai_alignment.txt) ## prerequisite - [AI alignment](../terminology/ai_alignment.txt) - [alignment target](../terminology/alignment_target.txt) - [training objective](../terminology/training_objective.txt) ## related - [reinforcement learning from human feedback](../terminology/reinforcement_learning_from_human_feedback.txt) ## Documents: about - [ベース目的関数とメサ目的関数は一致するとは限らない](../documents/hubinger_et_al.base-and-mesa-objectives-can-diverge.txt) [Corpus index](../index.txt)