--- aliases: en: - inner alignment problem concept_id: inner_alignment preferred: en: inner alignment ja: インナーアラインメント relations: relations: broader: - ai_alignment prerequisite: - ai_alignment - base_objective - mesa_objective - mesa_optimizer related: - objective_robustness - outer_alignment status: active --- # inner alignment The problem or property of ensuring that a mesa-optimizer's mesa-objective is robustly aligned with the base objective used by the base optimizer, rather than merely producing high-scoring behavior on the training distribution. In objective-focused terminology this is the mesa-optimizer special case of objective robustness; the exact scope of inner alignment varies across authors and settings. ## broader - [AI alignment](../terminology/ai_alignment.txt) ## prerequisite - [AI alignment](../terminology/ai_alignment.txt) - [base objective](../terminology/base_objective.txt) - [mesa-objective](../terminology/mesa_objective.txt) - [mesa-optimizer](../terminology/mesa_optimizer.txt) ## related - [objective robustness](../terminology/objective_robustness.txt) - [outer alignment](../terminology/outer_alignment.txt) ## Documents: about - [ベース目的関数とメサ目的関数は一致するとは限らない](../documents/hubinger_et_al.base-and-mesa-objectives-can-diverge.txt) - [観測された従順さだけでは、内部の目的や欺瞞の有無までは分からない](../documents/yudkowsky.observed-compliance-does-not-identify-hidden-objectives.txt) - [外部の目的関数は内部の目標を決定しない](../documents/yudkowsky.outer-objective-does-not-determine-inner-goal.txt) ## Documents: mentions - [能力はアラインメントより遠くまで汎化しうる](../documents/yudkowsky.capabilities-generalize-beyond-alignment.txt) - [能力と目的は別の軸である](../documents/yudkowsky.capability-motive-separation.txt) [Corpus index](../index.txt)