--- aliases: ja: - 目的頑健性 concept_id: objective_robustness preferred: en: objective robustness ja: 目的の頑健性 relations: relations: broader: - ai_alignment prerequisite: - base_objective - behavioral_objective related: - outer_alignment status: active --- # objective robustness The property that a learned system's behavioral objective remains aligned with the base objective it was trained or selected under, including when conditions differ from the training distribution. In the objective-focused inner-alignment framing, inner alignment is the mesa-optimizer special case of this broader concern; other generalization-focused framings use the term somewhat differently. ## broader - [AI alignment](../terminology/ai_alignment.txt) ## prerequisite - [base objective](../terminology/base_objective.txt) - [behavioral objective](../terminology/behavioral_objective.txt) ## related - [outer alignment](../terminology/outer_alignment.txt) ## Documents: about - [望ましい振る舞いは人間を重視する目標の証拠ではない](../documents/akira.behavioral-alignment-is-not-goal-alignment.txt) - [行動訓練で望ましく振る舞っても、意図した目標の頑健な汎化は保証されない](../documents/ngo_et_al.behavioral-training-does-not-establish-robust-goal-generalization.txt) [Corpus index](../index.txt)