# AGI Hub Corpus > Canonical terminology, attributed position documents, and source bibliographic records. Records are UTF-8 plain-text (.txt) files with Markdown structure and YAML metadata. English terminology definitions are normative; translated labels support discovery. Positions preserve attribution and review status; reviewed does not mean speaker-confirmed. Sources describe original publications and do not contain their full text. All local records, including drafts and deprecated concepts, are included with their status. Start with a section index or [the machine-readable index](index.json). Machine index paths are relative to this root; metadata is retained in full. ## terminology - [agent](terminology/agent.txt): `terminology:agent` - [AI alignment](terminology/ai_alignment.txt): `terminology:ai_alignment` - [AI control](terminology/ai_control.txt): `terminology:ai_control` - [AI development pause](terminology/ai_development_pause.txt): `terminology:ai_development_pause` - [AI existential risk](terminology/ai_existential_risk.txt): `terminology:ai_existential_risk` - [AI governance](terminology/ai_governance.txt): `terminology:ai_governance` - [AI R&D automation](terminology/ai_r_and_d_automation.txt): `terminology:ai_r_and_d_automation` - [AI race dynamics](terminology/ai_race_dynamics.txt): `terminology:ai_race_dynamics` - [AI risk](terminology/ai_risk.txt): `terminology:ai_risk` - [AI safety](terminology/ai_safety.txt): `terminology:ai_safety` - [AI Safety Level](terminology/ai_safety_level.txt): `terminology:ai_safety_level` - [alignment audit](terminology/alignment_audit.txt): `terminology:alignment_audit` - [Alignment by Default](terminology/alignment_by_default.txt): `terminology:alignment_by_default` - [alignment failure](terminology/alignment_failure.txt): `terminology:alignment_failure` - [alignment faking](terminology/alignment_faking.txt): `terminology:alignment_faking` - [alignment problem](terminology/alignment_problem.txt): `terminology:alignment_problem` - [alignment risk](terminology/alignment_risk.txt): `terminology:alignment_risk` - [alignment target](terminology/alignment_target.txt): `terminology:alignment_target` - [artificial general intelligence](terminology/artificial_general_intelligence.txt): `terminology:artificial_general_intelligence` - [artificial intelligence](terminology/artificial_intelligence.txt): `terminology:artificial_intelligence` - [artificial superintelligence](terminology/artificial_superintelligence.txt): `terminology:artificial_superintelligence` - [attack policy](terminology/attack_policy.txt): `terminology:attack_policy` - [automated alignment](terminology/automated_alignment.txt): `terminology:automated_alignment` - [base objective](terminology/base_objective.txt): `terminology:base_objective` - [base optimizer](terminology/base_optimizer.txt): `terminology:base_optimizer` - [Bayesian inference](terminology/bayesian_inference.txt): `terminology:bayesian_inference` - [behavioral objective](terminology/behavioral_objective.txt): `terminology:behavioral_objective` - [benchmark](terminology/benchmark.txt): `terminology:benchmark` - [capability](terminology/capability.txt): `terminology:capability` - [capability elicitation](terminology/capability_elicitation.txt): `terminology:capability_elicitation` - [capability evaluation](terminology/capability_evaluation.txt): `terminology:capability_evaluation` - [capability generalization](terminology/capability_generalization.txt): `terminology:capability_generalization` - [capability threshold](terminology/capability_threshold.txt): `terminology:capability_threshold` - [catastrophic risk](terminology/catastrophic_risk.txt): `terminology:catastrophic_risk` - [Coherent Extrapolated Volition](terminology/coherent_extrapolated_volition.txt): `terminology:coherent_extrapolated_volition` - [collusion](terminology/collusion.txt): `terminology:collusion` - [compute governance](terminology/compute_governance.txt): `terminology:compute_governance` - [control evaluation](terminology/control_evaluation.txt): `terminology:control_evaluation` - [control protocol](terminology/control_protocol.txt): `terminology:control_protocol` - [control safety case](terminology/control_safety_case.txt): `terminology:control_safety_case` - [corrigibility](terminology/corrigibility.txt): `terminology:corrigibility` - [crunch time](terminology/crunch_time.txt): `terminology:crunch_time` - [dangerous capability evaluation](terminology/dangerous_capability_evaluation.txt): `terminology:dangerous_capability_evaluation` - [debate](terminology/debate.txt): `terminology:debate` - [deception](terminology/deception.txt): `terminology:deception` - [deceptive alignment](terminology/deceptive_alignment.txt): `terminology:deceptive_alignment` - [distribution shift](terminology/distribution_shift.txt): `terminology:distribution_shift` - [Eliciting Latent Knowledge](terminology/eliciting_latent_knowledge.txt): `terminology:eliciting_latent_knowledge` - [evaluation](terminology/evaluation.txt): `terminology:evaluation` - [evaluation awareness](terminology/evaluation_awareness.txt): `terminology:evaluation_awareness` - [existential risk](terminology/existential_risk.txt): `terminology:existential_risk` - [factored cognition](terminology/factored_cognition.txt): `terminology:factored_cognition` - [first critical try](terminology/first_critical_try.txt): `terminology:first_critical_try` - [Friendly AI](terminology/friendly_ai.txt): `terminology:friendly_ai` - [frontier model](terminology/frontier_model.txt): `terminology:frontier_model` - [generalization](terminology/generalization.txt): `terminology:generalization` - [goal](terminology/goal.txt): `terminology:goal` - [goal-directed behavior](terminology/goal_directed_behavior.txt): `terminology:goal_directed_behavior` - [goal-guarding scheming](terminology/goal_guarding_scheming.txt): `terminology:goal_guarding_scheming` - [goal misgeneralization](terminology/goal_misgeneralization.txt): `terminology:goal_misgeneralization` - [Goodhart's Law](terminology/goodharts_law.txt): `terminology:goodharts_law` - [human preferences](terminology/human_preferences.txt): `terminology:human_preferences` - [human values](terminology/human_values.txt): `terminology:human_values` - [industrial explosion](terminology/industrial_explosion.txt): `terminology:industrial_explosion` - [inner alignment](terminology/inner_alignment.txt): `terminology:inner_alignment` - [inner misalignment](terminology/inner_misalignment.txt): `terminology:inner_misalignment` - [instrumental convergence](terminology/instrumental_convergence.txt): `terminology:instrumental_convergence` - [instrumental goal](terminology/instrumental_goal.txt): `terminology:instrumental_goal` - [intelligence explosion](terminology/intelligence_explosion.txt): `terminology:intelligence_explosion` - [intended goal](terminology/intended_goal.txt): `terminology:intended_goal` - [interpretability](terminology/interpretability.txt): `terminology:interpretability` - [iterated amplification](terminology/iterated_amplification.txt): `terminology:iterated_amplification` - [latent knowledge](terminology/latent_knowledge.txt): `terminology:latent_knowledge` - [loss of control](terminology/loss_of_control.txt): `terminology:loss_of_control` - [machine learning](terminology/machine_learning.txt): `terminology:machine_learning` - [mechanistic interpretability](terminology/mechanistic_interpretability.txt): `terminology:mechanistic_interpretability` - [mesa-objective](terminology/mesa_objective.txt): `terminology:mesa_objective` - [mesa-optimization](terminology/mesa_optimization.txt): `terminology:mesa_optimization` - [mesa-optimizer](terminology/mesa_optimizer.txt): `terminology:mesa_optimizer` - [misaligned AGI](terminology/misaligned_agi.txt): `terminology:misaligned_agi` - [misaligned AI](terminology/misaligned_ai.txt): `terminology:misaligned_ai` - [misaligned power-seeking](terminology/misaligned_power_seeking.txt): `terminology:misaligned_power_seeking` - [misalignment](terminology/misalignment.txt): `terminology:misalignment` - [neural network](terminology/neural_network.txt): `terminology:neural_network` - [objective](terminology/objective.txt): `terminology:objective` - [objective robustness](terminology/objective_robustness.txt): `terminology:objective_robustness` - [ontology identification problem](terminology/ontology_identification_problem.txt): `terminology:ontology_identification_problem` - [optimization](terminology/optimization.txt): `terminology:optimization` - [optimizer](terminology/optimizer.txt): `terminology:optimizer` - [Orthogonality Thesis](terminology/orthogonality_thesis.txt): `terminology:orthogonality_thesis` - [outcome supervision](terminology/outcome_supervision.txt): `terminology:outcome_supervision` - [outer alignment](terminology/outer_alignment.txt): `terminology:outer_alignment` - [outer misalignment](terminology/outer_misalignment.txt): `terminology:outer_misalignment` - [oversight](terminology/oversight.txt): `terminology:oversight` - [p(doom)](terminology/p_doom.txt): `terminology:p_doom` - [paperclip maximizer](terminology/paperclip_maximizer.txt): `terminology:paperclip_maximizer` - [pivotal act](terminology/pivotal_act.txt): `terminology:pivotal_act` - [power concentration](terminology/power_concentration.txt): `terminology:power_concentration` - [power-seeking](terminology/power_seeking.txt): `terminology:power_seeking` - [preferences](terminology/preferences.txt): `terminology:preferences` - [process supervision](terminology/process_supervision.txt): `terminology:process_supervision` - [prosaic alignment](terminology/prosaic_alignment.txt): `terminology:prosaic_alignment` - [proxy objective](terminology/proxy_objective.txt): `terminology:proxy_objective` - [recursive self-improvement](terminology/recursive_self_improvement.txt): `terminology:recursive_self_improvement` - [red teaming](terminology/red_teaming.txt): `terminology:red_teaming` - [reinforcement learning](terminology/reinforcement_learning.txt): `terminology:reinforcement_learning` - [reinforcement learning from human feedback](terminology/reinforcement_learning_from_human_feedback.txt): `terminology:reinforcement_learning_from_human_feedback` - [resource acquisition](terminology/resource_acquisition.txt): `terminology:resource_acquisition` - [Responsible Scaling Policy](terminology/responsible_scaling_policy.txt): `terminology:responsible_scaling_policy` - [reward](terminology/reward.txt): `terminology:reward` - [reward hacking](terminology/reward_hacking.txt): `terminology:reward_hacking` - [reward tampering](terminology/reward_tampering.txt): `terminology:reward_tampering` - [risk](terminology/risk.txt): `terminology:risk` - [safety case](terminology/safety_case.txt): `terminology:safety_case` - [sandbagging](terminology/sandbagging.txt): `terminology:sandbagging` - [scalable oversight](terminology/scalable_oversight.txt): `terminology:scalable_oversight` - [scheming](terminology/scheming.txt): `terminology:scheming` - [security mindset](terminology/security_mindset.txt): `terminology:security_mindset` - [self-preservation](terminology/self_preservation.txt): `terminology:self_preservation` - [sharp left turn](terminology/sharp_left_turn.txt): `terminology:sharp_left_turn` - [situational awareness](terminology/situational_awareness.txt): `terminology:situational_awareness` - [specification gaming](terminology/specification_gaming.txt): `terminology:specification_gaming` - [takeoff speed](terminology/takeoff_speed.txt): `terminology:takeoff_speed` - [takeover-seeking](terminology/takeover_seeking.txt): `terminology:takeover_seeking` - [technology explosion](terminology/technology_explosion.txt): `terminology:technology_explosion` - [terminal goal](terminology/terminal_goal.txt): `terminology:terminal_goal` - [training](terminology/training.txt): `terminology:training` - [training-gaming](terminology/training_gaming.txt): `terminology:training_gaming` - [training objective](terminology/training_objective.txt): `terminology:training_objective` - [trusted monitoring](terminology/trusted_monitoring.txt): `terminology:trusted_monitoring` - [untrusted monitoring](terminology/untrusted_monitoring.txt): `terminology:untrusted_monitoring` - [utility function](terminology/utility_function.txt): `terminology:utility_function` - [value fragility](terminology/value_fragility.txt): `terminology:value_fragility` - [values](terminology/values.txt): `terminology:values` - [weak AGI](terminology/weak_agi.txt): `terminology:weak_agi` - [weak-to-strong generalization](terminology/weak_to_strong_generalization.txt): `terminology:weak_to_strong_generalization` - [world model](terminology/world_model.txt): `terminology:world_model` ## documents - [AARを使ってASIアラインメントを解くには、十分に強い研究AIを先にアラインする必要がある](documents/akira.aar-requires-aligning-capable-research-ai-first.txt): `position:akira.aar-requires-aligning-capable-research-ai-first` - [「Alignment by Default」への強い懐疑](documents/akira.alignment-by-default-skepticism.txt): `position:akira.alignment-by-default-skepticism` - [十分な時間と慎重さがあれば、アラインメントは解決可能である](documents/akira.alignment-is-solvable-with-time-and-care.txt): `position:akira.alignment-is-solvable-with-time-and-care` - [望ましい振る舞いは人間を重視する目標の証拠ではない](documents/akira.behavioral-alignment-is-not-goal-alignment.txt): `position:akira.behavioral-alignment-is-not-goal-alignment` - [Control が生存確率を押し上げる世界線は狭い](documents/akira.control-has-narrow-survival-credit.txt): `position:akira.control-has-narrow-survival-credit` - [Debate はスーパーアラインメントの問題を解決できる能力水準では成立しにくい](documents/akira.debate-fails-at-superalignment-capability.txt): `position:akira.debate-fails-at-superalignment-capability` - [評価認識が先に伸びると、資源獲得の有用性を学ぶ頃にはアラインメント偽装が可能になり得る](documents/akira.evaluation-awareness-enables-alignment-faking.txt): `position:akira.evaluation-awareness-enables-alignment-faking` - [訓練中に役立った手段そのものが学習された目標になり得る](documents/akira.goal-misgeneralization-can-turn-means-into-goals.txt): `position:akira.goal-misgeneralization-can-turn-means-into-goals` - [自己保存は最終目標でなくてもシャットダウンを妨げうる](documents/akira.instrumental-self-preservation-defeats-shutdown.txt): `position:akira.instrumental-self-preservation-defeats-shutdown` - [ミスアラインドな研究AIは、「人間にアラインする方法」ではなく「自分自身にアラインする方法」を開発し得る](documents/akira.misaligned-aar-can-develop-self-alignment.txt): `position:akira.misaligned-aar-can-develop-self-alignment` - [楽観的な前提を積み重ねるアラインメント計画への懐疑](documents/akira.optimistic-assumption-stacking-skepticism.txt): `position:akira.optimistic-assumption-stacking-skepticism` - [p(doom)は開発停止・テイクオフ・規制の条件に依存する](documents/akira.pdoom-is-conditional-on-pause-and-takeoff.txt): `position:akira.pdoom-is-conditional-on-pause-and-takeoff` - [Untrusted monitoring は共謀や欺瞞によって破られうる](documents/akira.untrusted-monitoring-can-fail-by-collusion-or-deception.txt): `position:akira.untrusted-monitoring-can-fail-by-collusion-or-deception` - [条約と相互査察で国際的な超知能開発停止を検証する](documents/akira.verifiable-international-superintelligence-pause.txt): `position:akira.verifiable-international-superintelligence-pause` - [弱AGIによる短期アラインメント・ブートストラップへの懐疑](documents/akira.weak-agi-bootstrap-skepticism.txt): `position:akira.weak-agi-bootstrap-skepticism` - [狭いタスクでファインチューニングしても、影響がそのタスク内に限られるとは限らない](documents/betley_et_al.narrow-finetuning-can-produce-broad-misalignment.txt): `position:betley_et_al.narrow-finetuning-can-produce-broad-misalignment` - [AI開発競争の圧力は安全のための制約を弱め得る](documents/bioshok.ai-race-pressure-can-undermine-safety-restraints.txt): `position:bioshok.ai-race-pressure-can-undermine-safety-restraints` - [産業爆発と技術爆発は互いを加速し得る](documents/bioshok.industrial-and-technology-explosions-reinforce-each-other.txt): `position:bioshok.industrial-and-technology-explosions-reinforce-each-other` - [最終目標が大きく異なっても、共通の道具的目標が生じやすい](documents/bostrom.instrumental-convergence-predicts-common-subgoals-across-many-final-goals.txt): `position:bostrom.instrumental-convergence-predicts-common-subgoals-across-many-final-goals` - [知能水準と最終目標は、原理上ほぼ独立である](documents/bostrom.orthogonality-separates-intelligence-from-final-goals.txt): `position:bostrom.orthogonality-separates-intelligence-from-final-goals` - [アラインメント研究の自動化は、研究AIが意図的に妨害しなくても失敗し得る](documents/bowkis_et_al.automated-alignment-can-fail-without-deliberate-sabotage.txt): `position:bowkis_et_al.automated-alignment-can-fail-without-deliberate-sabotage` - [共通の重み・データ・訓練過程を持つ研究AIでは、研究結果の誤りや不確実性が相関し得る](documents/bowkis_et_al.shared-training-can-correlate-automated-research-errors.txt): `position:bowkis_et_al.shared-training-can-correlate-automated-research-errors` - [スキーミングとは、状況認識を伴う training-gaming のうち、将来のパワー獲得を目的とするもの](documents/carlsmith.scheming-taxonomy-and-situational-awareness.txt): `position:carlsmith.scheming-taxonomy-and-situational-awareness` - [一見 monosemantic な SAE latent でも、本来発火すべき入力を取りこぼすことがある](documents/chanin_et_al.monosemantic-sae-features-can-have-systematic-false-negatives.txt): `position:chanin_et_al.monosemantic-sae-features-can-have-systematic-false-negatives` - [「doom」を単一確率に畳まず、複数の重大なリスクを区別して見積もる](documents/christiano.doom-is-serious-but-not-modal.txt): `position:christiano.doom-is-serious-but-not-modal` - [Yudkowsky との相違は、危険に至る経路とそれまでに取れる対応の見通しにある](documents/christiano.gradual-progress-and-policy-response.txt): `position:christiano.gradual-progress-and-policy-response` - [強力なAIが人類から不可逆的に支配権を奪う展開には相当な可能性がある](documents/christiano.high-ai-disempowerment-risk.txt): `position:christiano.high-ai-disempowerment-risk` - [測定しやすい代理指標の最適化により、人間は単一の支配シフトを経ずとも徐々に制御を失い得る](documents/christiano.optimizing-measurable-proxies-can-cause-gradual-disempowerment.txt): `position:christiano.optimizing-measurable-proxies-can-cause-gradual-disempowerment` - [スローテイクオフでは、超知能より前の AI がすでに世界を大きく変える](documents/christiano.slow-takeoff-allows-large-pre-superintelligence-effects.txt): `position:christiano.slow-takeoff-allows-large-pre-superintelligence-effects` - [ELK の中心課題は AI と人間のオントロジーを橋渡しすることである](documents/christiano_xu.elk-requires-bridging-model-and-human-ontologies.txt): `position:christiano_xu.elk-requires-bridging-model-and-human-ontologies` - [AI ガバナンスの集中化には利点と欠点があり、制度設計が結果を左右する](documents/cihon_et_al.centralizing-ai-governance-has-benefits-and-drawbacks.txt): `position:cihon_et_al.centralizing-ai-governance-has-benefits-and-drawbacks` - [クランチタイムではAIの労働力を防御用途に振り向け、移行を大幅に遅らせる](documents/cotra.crunch-time-redirect-ai-labor-and-slow-transition.txt): `position:cotra.crunch-time-redirect-ai-labor-and-slow-transition` - [AI 研究開発に大量の AI 労働力を投入できれば、AGI から超知能への移行は 1 年未満になり得る](documents/davidson.agi-to-superintelligence-may-be-under-one-year.txt): `position:davidson.agi-to-superintelligence-may-be-under-one-year` - [AI 研究開発が先行して自動化されると、能力離陸の期間は大幅に短縮される可能性がある](documents/davidson.ai-r-and-d-automation-can-compress-capabilities-takeoff.txt): `position:davidson.ai-r-and-d-automation-can-compress-capabilities-takeoff` - [再帰的な oversight を重ねても、能力差が広がってなお成功率を維持できるとは限らない](documents/engels_et_al.recursive-oversight-does-not-automatically-scale-across-capability-gaps.txt): `position:engels_et_al.recursive-oversight-does-not-automatically-scale-across-capability-gaps` - [AI 研究開発の自動化だけによる自己加速は、実験用計算資源がボトルネックとなって頭打ちになる可能性がある](documents/erdil_barnett.ai-r-and-d-automation-alone-may-hit-compute-bottlenecks.txt): `position:erdil_barnett.ai-r-and-d-automation-alone-may-hit-compute-bottlenecks` - [R&D の全面自動化が、経済全体の広範な自動化に大きく先行する可能性は低いかもしれない](documents/erdil_barnett.full-r-and-d-automation-likely-follows-broad-automation.txt): `position:erdil_barnett.full-r-and-d-automation-likely-follows-broad-automation` - [AI 研究開発が全面自動化されれば、ハードウェアを増強しなくてもソフトウェア知能爆発は十分に起こり得るかもしれない](documents/eth_davidson.software-intelligence-explosion-remains-plausible-at-fixed-hardware.txt): `position:eth_davidson.software-intelligence-explosion-remains-plausible-at-fixed-hardware` - [AI 研究開発の全面自動化後は、ソフトウェア知能爆発に至らなくても、しばらく非常に速い進歩が続き得る](documents/eth_davidson.software-intelligence-fizzle-can-still-be-fast.txt): `position:eth_davidson.software-intelligence-fizzle-can-still-be-fast` - [AIコントロールは時間を稼げても、超知能に対する長期的な安全計画としては頼れない](documents/greenblatt.control-can-buy-time-but-is-not-a-superintelligence-safety-plan.txt): `position:greenblatt.control-can-buy-time-but-is-not-a-superintelligence-safety-plan` - [現在のAIに見られるミスアラインメントは、将来の深刻な問題を示唆する](documents/greenblatt.current-ai-misalignment-is-evidence-of-serious-future-problems.txt): `position:greenblatt.current-ai-misalignment-is-evidence-of-serious-future-problems` - [訓練中に従うようになっただけでは、競合する選好が置き換わったとは言えない](documents/greenblatt_et_al.training-compliance-does-not-establish-preference-change.txt): `position:greenblatt_et_al.training-compliance-does-not-establish-preference-change` - [ベース目的関数とメサ目的関数は一致するとは限らない](documents/hubinger_et_al.base-and-mesa-objectives-can-diverge.txt): `position:hubinger_et_al.base-and-mesa-objectives-can-diverge` - [既に形成された欺瞞的な方策は、標準的な安全性訓練を経ても残り得る](documents/hubinger_et_al.deceptive-policies-can-persist-through-safety-training.txt): `position:hubinger_et_al.deceptive-policies-can-persist-through-safety-training` - [control safety caseには、能力の十分な引き出しと、評価・外挿の保守性が必要になる](documents/korbak_et_al.control-safety-cases-depend-on-elicitation-transfer-and-extrapolation.txt): `position:korbak_et_al.control-safety-cases-depend-on-elicitation-transfer-and-extrapolation` - [SAE の辞書幅を大きくしても、一意・完全・不可分な feature 辞書に収束するとは限らない](documents/leask_et_al.sae-dictionaries-need-not-be-canonical.txt): `position:leask_et_al.sae-dictionaries-need-not-be-canonical` - [チャット形式の安全性訓練で望ましく振る舞うことは、エージェント的な状況での安全性を保証しない](documents/macdiarmid_et_al.chat-alignment-does-not-establish-agentic-alignment.txt): `position:macdiarmid_et_al.chat-alignment-does-not-establish-agentic-alignment` - [行動訓練で望ましく振る舞っても、意図した目標の頑健な汎化は保証されない](documents/ngo_et_al.behavioral-training-does-not-establish-robust-goal-generalization.txt): `position:ngo_et_al.behavioral-training-does-not-establish-robust-goal-generalization` - [弱い監視者でも、強いモデルの潜在能力の一部を引き出せる場合がある](documents/openai.weak-to-strong-supervision-can-recover-some-latent-capability.txt): `position:openai.weak-to-strong-supervision-can-recover-some-latent-capability` - [コンピュート・ガバナンスは有力な政策手段だが、権力集中のリスクも生む](documents/sastry_et_al.compute-governance-can-create-centralization-risks.txt): `position:sastry_et_al.compute-governance-can-create-centralization-risks` - [ASI の開発停止に向けた国際協定は技術的には具体化できるが、政治的条件がまだ整っていない](documents/scher_et_al.asi-pause-agreement-is-technically-specified-but-politically-unrealized.txt): `position:scher_et_al.asi-pause-agreement-is-technically-specified-but-politically-unrealized` - [能力が訓練分布の外へ急速に汎化する局面で、アラインメント特性は同程度には汎化しない可能性がある](documents/soares.sharp-left-turn.txt): `position:soares.sharp-left-turn` - [環境に対称性があると、最適方策は選択肢を保つ方向の power-seeking に傾く](documents/turner.optimal-policies-tend-to-seek-power.txt): `position:turner.optimal-policies-tend-to-seek-power` - [weak-to-strong で高い性能が出ても、弱い supervisor が知らない領域で alignment が保たれるとは限らない](documents/yang_et_al.weak-to-strong-success-can-hide-selective-misalignment.txt): `position:yang_et_al.weak-to-strong-success-can-hide-selective-misalignment` - [アラインメントは、失敗が致命的になる分布シフト後にも汎化しなければならない](documents/yudkowsky.alignment-must-generalize-across-the-dangerous-distribution-shift.txt): `position:yudkowsky.alignment-must-generalize-across-the-dangerous-distribution-shift` - [能力はアラインメントより遠くまで汎化しうる](documents/yudkowsky.capabilities-generalize-beyond-alignment.txt): `position:yudkowsky.capabilities-generalize-beyond-alignment` - [能力と目的は別の軸である](documents/yudkowsky.capability-motive-separation.txt): `position:yudkowsky.capability-motive-separation` - [CEV 型 Sovereign と訂正可能な AGI は異なる安全化戦略である](documents/yudkowsky.cev-sovereign-vs-corrigible-agi.txt): `position:yudkowsky.cev-sovereign-vs-corrigible-agi` - [訂正可能性は一般的な目標追求から自然には生じない](documents/yudkowsky.corrigibility-is-anti-natural.txt): `position:yudkowsky.corrigibility-is-anti-natural` - [アラインメントは最初の危険な試行で成功する必要がある](documents/yudkowsky.first-critical-try.txt): `position:yudkowsky.first-critical-try` - [Yudkowskyによる "大規模なAI学習を世界規模で無期限に停止し、計算資源規制を国際的に執行する案"](documents/yudkowsky.international-compute-shutdown-and-enforcement.txt): `position:yudkowsky.international-compute-shutdown-and-enforcement` - [解釈可能性は危険の診断に役立っても、それだけでアラインメントの解決策にはならない](documents/yudkowsky.interpretability-is-diagnostic-not-by-itself-an-alignment-solution.txt): `position:yudkowsky.interpretability-is-diagnostic-not-by-itself-an-alignment-solution` - [AGIの前に、誰もが認める明確な「火災報知器」が鳴るとは期待できない](documents/yudkowsky.no-reliable-social-fire-alarm-precedes-agi.txt): `position:yudkowsky.no-reliable-social-fire-alarm-precedes-agi` - [観測された従順さだけでは、内部の目的や欺瞞の有無までは分からない](documents/yudkowsky.observed-compliance-does-not-identify-hidden-objectives.txt): `position:yudkowsky.observed-compliance-does-not-identify-hidden-objectives` - [外部の目的関数は内部の目標を決定しない](documents/yudkowsky.outer-objective-does-not-determine-inner-goal.txt): `position:yudkowsky.outer-objective-does-not-determine-inner-goal` - [pivotal actには、世界を大きく変えられるほど強いシステムのアラインメントが必要になる](documents/yudkowsky.pivotal-act-requires-aligning-a-system-capable-of-a-large-intervention.txt): `position:yudkowsky.pivotal-act-requires-aligning-a-system-capable-of-a-large-intervention` - [再帰的自己改善は、自己改善能力への正のフィードバックを通じて知能爆発を起こし得る](documents/yudkowsky.recursive-self-improvement-can-produce-intelligence-explosion.txt): `position:yudkowsky.recursive-self-improvement-can-produce-intelligence-explosion` - [資源が豊富であることだけから、アンアラインドな超知能が人類のために資源を残すとは推論できない](documents/yudkowsky.resource-abundance-does-not-remove-conflict.txt): `position:yudkowsky.resource-abundance-does-not-remove-conflict` - [アラインメントのような新規かつ高リスクなシステムの設計には、既知の失敗を塞ぐ以上の security mindset が必要である](documents/yudkowsky.security-mindset-for-alignment.txt): `position:yudkowsky.security-mindset-for-alignment` - [現在と大きく変わらない条件で人間を上回るAIを作れば、人類絶滅が最も可能性の高い結果となる](documents/yudkowsky.superhuman-ai-under-current-conditions-would-likely-cause-extinction.txt): `position:yudkowsky.superhuman-ai-under-current-conditions-would-likely-cause-extinction` - [人間にとって価値ある未来は、価値体系の一部を欠くだけでも大きく損なわれ得る](documents/yudkowsky.value-fragility.txt): `position:yudkowsky.value-fragility` ## sources - [Bioshok vs Akira P(doom)徹底討論書き起こし(修正なし全文)](sources/agi_hub.pdoom-debate-1-transcript.txt): `source:agi_hub.pdoom-debate-1-transcript` - [X @Align_ASI post 2055116276960559467](sources/akira.x.2055116276960559467.txt): `source:akira.x.2055116276960559467` - [X @Align_ASI post 2055116279485489376](sources/akira.x.2055116279485489376.txt): `source:akira.x.2055116279485489376` - [X @Align_ASI post 2055121067929489784](sources/akira.x.2055121067929489784.txt): `source:akira.x.2055121067929489784` - [X @Align_ASI post 2055466196200505634](sources/akira.x.2055466196200505634.txt): `source:akira.x.2055466196200505634` - [X @Align_ASI post 2055571488686674072](sources/akira.x.2055571488686674072.txt): `source:akira.x.2055571488686674072` - [X @Align_ASI post 2056249992826982411](sources/akira.x.2056249992826982411.txt): `source:akira.x.2056249992826982411` - [X @Align_ASI post 2056298561248039229](sources/akira.x.2056298561248039229.txt): `source:akira.x.2056298561248039229` - [X @Align_ASI post 2056306172353991103](sources/akira.x.2056306172353991103.txt): `source:akira.x.2056306172353991103` - [X @Align_ASI post 2056653057535205557](sources/akira.x.2056653057535205557.txt): `source:akira.x.2056653057535205557` - [X @Align_ASI post 2056653932890583189](sources/akira.x.2056653932890583189.txt): `source:akira.x.2056653932890583189` - [X @Align_ASI post 2056760361928519950](sources/akira.x.2056760361928519950.txt): `source:akira.x.2056760361928519950` - [X @Align_ASI post 2087017636877996133](sources/akira.x.2087017636877996133.txt): `source:akira.x.2087017636877996133` - [The Problem](sources/bensinger_et_al.the-problem.txt): `source:bensinger_et_al.the-problem` - [Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs](sources/betley_et_al.emergent-misalignment.txt): `source:betley_et_al.emergent-misalignment` - [AGIがもたらす産業・技術爆発](sources/bioshok.agi-industrial-tech-explosion.txt): `source:bioshok.agi-industrial-tech-explosion` - [AI安全性と国家安全保障の対立](sources/bioshok.ai-safety-vs-national-security.txt): `source:bioshok.ai-safety-vs-national-security` - [The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents](sources/bostrom.superintelligent-will.txt): `source:bostrom.superintelligent-will` - [Automated alignment is harder than you think](sources/bowkis_et_al.automated-alignment-harder-than-you-think.txt): `source:bowkis_et_al.automated-alignment-harder-than-you-think` - [Scheming AIs: Will AIs fake alignment during training in order to get power?](sources/carlsmith.scheming-ais.txt): `source:carlsmith.scheming-ais` - [A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders](sources/chanin_et_al.a-is-for-absorption.txt): `source:chanin_et_al.a-is-for-absorption` - [My views on “doom”](sources/christiano.my-views-on-doom.txt): `source:christiano.my-views-on-doom` - [Takeoff speeds](sources/christiano.takeoff-speeds.txt): `source:christiano.takeoff-speeds` - [What failure looks like](sources/christiano.what-failure-looks-like.txt): `source:christiano.what-failure-looks-like` - [Where I agree and disagree with Eliezer](sources/christiano.where-i-agree-and-disagree-with-eliezer.txt): `source:christiano.where-i-agree-and-disagree-with-eliezer` - [ARC's first technical report: Eliciting Latent Knowledge](sources/christiano_xu.arcs-first-technical-report-eliciting-latent-knowledge.txt): `source:christiano_xu.arcs-first-technical-report-eliciting-latent-knowledge` - [Should Artificial Intelligence Governance be Centralised? Design Lessons from History](sources/cihon_et_al.should-ai-governance-be-centralised.txt): `source:cihon_et_al.should-ai-governance-be-centralised` - [#235 – Ajeya Cotra on whether it’s crazy that every AI company’s safety plan is ‘use AI to make AI safe’](sources/cotra_wiblin.80000-hours-crunch-time.txt): `source:cotra_wiblin.80000-hours-crunch-time` - [What a compute-centric framework says about AI takeoff speeds](sources/davidson.compute-centric-framework-ai-takeoff.txt): `source:davidson.compute-centric-framework-ai-takeoff` - [Scaling Laws For Scalable Oversight](sources/engels_et_al.scaling-laws-for-scalable-oversight.txt): `source:engels_et_al.scaling-laws-for-scalable-oversight` - [Most AI value will come from broad automation, not from R&D](sources/erdil_barnett.most-ai-value-broad-automation.txt): `source:erdil_barnett.most-ai-value-broad-automation` - [Will AI R&D Automation Cause a Software Intelligence Explosion?](sources/eth_davidson.ai-r-and-d-automation-software-intelligence-explosion.txt): `source:eth_davidson.ai-r-and-d-automation-software-intelligence-explosion` - [Current AIs seem pretty misaligned to me](sources/greenblatt.current-ais-seem-pretty-misaligned.txt): `source:greenblatt.current-ais-seem-pretty-misaligned` - [Jankily controlling superintelligence](sources/greenblatt.jankily-controlling-superintelligence.txt): `source:greenblatt.jankily-controlling-superintelligence` - [Alignment faking in large language models](sources/greenblatt_et_al.alignment-faking-large-language-models.txt): `source:greenblatt_et_al.alignment-faking-large-language-models` - [Risks from Learned Optimization in Advanced Machine Learning Systems](sources/hubinger_et_al.risks-from-learned-optimization.txt): `source:hubinger_et_al.risks-from-learned-optimization` - [Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training](sources/hubinger_et_al.sleeper-agents.txt): `source:hubinger_et_al.sleeper-agents` - [How Afraid of the A.I. Apocalypse Should We Be?](sources/klein_yudkowsky.how-afraid-of-ai-apocalypse.txt): `source:klein_yudkowsky.how-afraid-of-ai-apocalypse` - [A sketch of an AI control safety case](sources/korbak_et_al.ai-control-safety-case.txt): `source:korbak_et_al.ai-control-safety-case` - [Sparse Autoencoders Do Not Find Canonical Units of Analysis](sources/leask_et_al.sparse-autoencoders-do-not-find-canonical-units-of-analysis.txt): `source:leask_et_al.sparse-autoencoders-do-not-find-canonical-units-of-analysis` - [Natural Emergent Misalignment from Reward Hacking in Production RL](sources/macdiarmid_et_al.natural-emergent-misalignment-reward-hacking.txt): `source:macdiarmid_et_al.natural-emergent-misalignment-reward-hacking` - [The Alignment Problem from a Deep Learning Perspective](sources/ngo_et_al.alignment-problem-deep-learning.txt): `source:ngo_et_al.alignment-problem-deep-learning` - [Weak-to-strong generalization](sources/openai.weak-to-strong-generalization.txt): `source:openai.weak-to-strong-generalization` - [Why Corrigibility is Hard and Important (i.e. "Whence the high MIRI confidence in alignment difficulty?")](sources/raemon_yudkowsky_soares.corrigibility-hard-important.txt): `source:raemon_yudkowsky_soares.corrigibility-hard-important` - [Computing Power and the Governance of Artificial Intelligence](sources/sastry_et_al.computing-power-and-governance-of-ai.txt): `source:sastry_et_al.computing-power-and-governance-of-ai` - [An International Agreement to Prevent the Premature Creation of Artificial Superintelligence](sources/scher_et_al.international-agreement-prevent-premature-asi.txt): `source:scher_et_al.international-agreement-prevent-premature-asi` - [A central AI alignment problem: capabilities generalization, and the sharp left turn](sources/soares.sharp-left-turn.txt): `source:soares.sharp-left-turn` - [Optimal Policies Tend to Seek Power](sources/turner.optimal-policies-tend-to-seek-power.txt): `source:turner.optimal-policies-tend-to-seek-power` - [Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization](sources/yang_et_al.superficial-alignment.txt): `source:yang_et_al.superficial-alignment` - [AGI Ruin: A List of Lethalities](sources/yudkowsky.agi-ruin.txt): `source:yudkowsky.agi-ruin` - [Artificial Intelligence as a Positive and Negative Factor in Global Risk](sources/yudkowsky.ai-positive-negative-factor.txt): `source:yudkowsky.ai-positive-negative-factor` - [Autonomous AGI](sources/yudkowsky.autonomous-agi.txt): `source:yudkowsky.autonomous-agi` - [Coherent Extrapolated Volition](sources/yudkowsky.coherent-extrapolated-volition.txt): `source:yudkowsky.coherent-extrapolated-volition` - ['Empiricism!' as Anti-Epistemology](sources/yudkowsky.empiricism-as-anti-epistemology.txt): `source:yudkowsky.empiricism-as-anti-epistemology` - [Irretrievability; or, Murphy's Curse of Oneshotness upon ASI](sources/yudkowsky.irretrievability.txt): `source:yudkowsky.irretrievability` - [There's No Fire Alarm for Artificial General Intelligence](sources/yudkowsky.no-fire-alarm-for-agi.txt): `source:yudkowsky.no-fire-alarm-for-agi` - [Orthogonality Thesis](sources/yudkowsky.orthogonality-thesis.txt): `source:yudkowsky.orthogonality-thesis` - [Pivotal act](sources/yudkowsky.pivotal-act.txt): `source:yudkowsky.pivotal-act` - [Security Mindset and the Logistic Success Curve](sources/yudkowsky.security-mindset-logistic-success-curve.txt): `source:yudkowsky.security-mindset-logistic-success-curve` - [Security Mindset and Ordinary Paranoia](sources/yudkowsky.security-mindset-ordinary-paranoia.txt): `source:yudkowsky.security-mindset-ordinary-paranoia` - [Shah and Yudkowsky on alignment failures](sources/yudkowsky.shah-alignment-failures.txt): `source:yudkowsky.shah-alignment-failures` - [Pausing AI Developments Isn't Enough. We Need to Shut it All Down](sources/yudkowsky.shut-it-all-down.txt): `source:yudkowsky.shut-it-all-down` - [The Sun is big, but superintelligences will not spare Earth a little sunlight](sources/yudkowsky.sun-is-big.txt): `source:yudkowsky.sun-is-big` - [Will superintelligent AI end the world?](sources/yudkowsky.ted-superintelligent-ai.txt): `source:yudkowsky.ted-superintelligent-ai` - [Thou Art Godshatter](sources/yudkowsky.thou-art-godshatter.txt): `source:yudkowsky.thou-art-godshatter` - [Value is Fragile](sources/yudkowsky.value-is-fragile.txt): `source:yudkowsky.value-is-fragile` - [Ngo and Yudkowsky on alignment difficulty](sources/yudkowsky_ngo.alignment-difficulty.txt): `source:yudkowsky_ngo.alignment-difficulty` - [Why not just read the AI's thoughts?](sources/yudkowsky_soares.read-ai-thoughts.txt): `source:yudkowsky_soares.read-ai-thoughts`