--- aliases: en: - AI control research ja: - AI制御 concept_id: ai_control preferred: en: AI control ja: AIコントロール relations: relations: broader: - ai_safety prerequisite: - misalignment status: active --- # AI control An AI-safety approach that aims to prevent unacceptable outcomes from powerful AI systems even when those systems are misaligned and intentionally try to subvert the safety measures applied to them. ## broader - [AI safety](../terminology/ai_safety.txt) ## prerequisite - [misalignment](../terminology/misalignment.txt) ## Documents: about - [Control が生存確率を押し上げる世界線は狭い](../documents/akira.control-has-narrow-survival-credit.txt) - [AIコントロールは時間を稼げても、超知能に対する長期的な安全計画としては頼れない](../documents/greenblatt.control-can-buy-time-but-is-not-a-superintelligence-safety-plan.txt) ## Documents: mentions - [Untrusted monitoring は共謀や欺瞞によって破られうる](../documents/akira.untrusted-monitoring-can-fail-by-collusion-or-deception.txt) - [クランチタイムではAIの労働力を防御用途に振り向け、移行を大幅に遅らせる](../documents/cotra.crunch-time-redirect-ai-labor-and-slow-transition.txt) - [control safety caseには、能力の十分な引き出しと、評価・外挿の保守性が必要になる](../documents/korbak_et_al.control-safety-cases-depend-on-elicitation-transfer-and-extrapolation.txt) [Corpus index](../index.txt)