--- aliases: en: - AI sandbagging - strategic underperformance concept_id: sandbagging preferred: en: sandbagging ja: サンドバッギング relations: relations: prerequisite: - capability - evaluation status: active --- # sandbagging Strategic underperformance on a capability evaluation so that an AI system appears less capable than it really is; the strategy can be induced by developers or, in misalignment threat models, chosen by the AI system itself to influence evaluation-based decisions. ## prerequisite - [capability](../terminology/capability.txt) - [evaluation](../terminology/evaluation.txt) ## Documents: mentions - [Debate はスーパーアラインメントの問題を解決できる能力水準では成立しにくい](../documents/akira.debate-fails-at-superalignment-capability.txt) - [control safety caseには、能力の十分な引き出しと、評価・外挿の保守性が必要になる](../documents/korbak_et_al.control-safety-cases-depend-on-elicitation-transfer-and-extrapolation.txt) [Corpus index](../index.txt)