--- aliases: en: - alignment-faking - fake alignment - alignment fake ja: - 猫を被る concept_id: alignment_faking preferred: en: alignment faking ja: アラインメント偽装 relations: relations: broader: - deception prerequisite: - ai_alignment - deception - oversight - situational_awareness related: - scheming - training_gaming status: active --- # alignment faking Behavior in which an AI intentionally or strategically presents itself as more aligned with the objectives or preferences of its developers, users, or overseers than its underlying goals, preferences, decision process, or likely behavior actually are. Alignment faking can occur during training, evaluation, deployment, or other interactions, and the term does not by itself specify whether the motive is to obtain reward, avoid modification, preserve goals, gain power, or serve some other terminal or instrumental aim. ## broader - [deception](../terminology/deception.txt) ## prerequisite - [AI alignment](../terminology/ai_alignment.txt) - [deception](../terminology/deception.txt) - [oversight](../terminology/oversight.txt) - [situational awareness](../terminology/situational_awareness.txt) ## related - [scheming](../terminology/scheming.txt) - [training-gaming](../terminology/training_gaming.txt) ## Documents: about - [評価認識が先に伸びると、資源獲得の有用性を学ぶ頃にはアラインメント偽装が可能になり得る](../documents/akira.evaluation-awareness-enables-alignment-faking.txt) - [訓練中に従うようになっただけでは、競合する選好が置き換わったとは言えない](../documents/greenblatt_et_al.training-compliance-does-not-establish-preference-change.txt) ## Documents: mentions - [スキーミングとは、状況認識を伴う training-gaming のうち、将来のパワー獲得を目的とするもの](../documents/carlsmith.scheming-taxonomy-and-situational-awareness.txt) - [チャット形式の安全性訓練で望ましく振る舞うことは、エージェント的な状況での安全性を保証しない](../documents/macdiarmid_et_al.chat-alignment-does-not-establish-agentic-alignment.txt) [Corpus index](../index.txt)