--- aliases: en: - AI deception - strategic deception ja: - デセプション concept_id: deception preferred: en: deception ja: 欺瞞 status: active --- # deception Intentional or strategic behavior that causes another agent or evaluator to form a materially false or misleading belief; in AI safety, deception is broader than deceptive alignment and can occur for many reasons unrelated to the training process. ## Documents: about - [観測された従順さだけでは、内部の目的や欺瞞の有無までは分からない](../documents/yudkowsky.observed-compliance-does-not-identify-hidden-objectives.txt) ## Documents: mentions - [Untrusted monitoring は共謀や欺瞞によって破られうる](../documents/akira.untrusted-monitoring-can-fail-by-collusion-or-deception.txt) - [訓練中に従うようになっただけでは、競合する選好が置き換わったとは言えない](../documents/greenblatt_et_al.training-compliance-does-not-establish-preference-change.txt) - [既に形成された欺瞞的な方策は、標準的な安全性訓練を経ても残り得る](../documents/hubinger_et_al.deceptive-policies-can-persist-through-safety-training.txt) - [解釈可能性は危険の診断に役立っても、それだけでアラインメントの解決策にはならない](../documents/yudkowsky.interpretability-is-diagnostic-not-by-itself-an-alignment-solution.txt) [Corpus index](../index.txt)