AGI Hub Corpus
Raw UTF-8 Markdown and machine-readable metadata.
Corpus guide
JSON index
LLM guide
terminology
agent
AI alignment
AI control
AI development pause
AI existential risk
AI governance
AI R&D automation
AI race dynamics
AI risk
AI safety
AI Safety Level
alignment audit
Alignment by Default
alignment failure
alignment faking
alignment problem
alignment risk
alignment target
artificial general intelligence
artificial intelligence
artificial superintelligence
attack policy
automated alignment
base objective
base optimizer
Bayesian inference
behavioral objective
benchmark
capability
capability elicitation
capability evaluation
capability generalization
capability threshold
catastrophic risk
Coherent Extrapolated Volition
collusion
compute governance
control evaluation
control protocol
control safety case
corrigibility
crunch time
dangerous capability evaluation
debate
deception
deceptive alignment
distribution shift
Eliciting Latent Knowledge
evaluation
evaluation awareness
existential risk
factored cognition
first critical try
Friendly AI
frontier model
generalization
goal
goal-directed behavior
goal-guarding scheming
goal misgeneralization
Goodhart's Law
human preferences
human values
industrial explosion
inner alignment
inner misalignment
instrumental convergence
instrumental goal
intelligence explosion
intended goal
interpretability
iterated amplification
latent knowledge
loss of control
machine learning
mechanistic interpretability
mesa-objective
mesa-optimization
mesa-optimizer
misaligned AGI
misaligned AI
misaligned power-seeking
misalignment
neural network
objective
objective robustness
ontology identification problem
optimization
optimizer
Orthogonality Thesis
outcome supervision
outer alignment
outer misalignment
oversight
p(doom)
paperclip maximizer
pivotal act
power concentration
power-seeking
preferences
process supervision
prosaic alignment
proxy objective
recursive self-improvement
red teaming
reinforcement learning
reinforcement learning from human feedback
resource acquisition
Responsible Scaling Policy
reward
reward hacking
reward tampering
risk
safety case
sandbagging
scalable oversight
scheming
security mindset
self-preservation
sharp left turn
situational awareness
specification gaming
takeoff speed
takeover-seeking
technology explosion
terminal goal
training
training-gaming
training objective
trusted monitoring
untrusted monitoring
utility function
value fragility
values
weak AGI
weak-to-strong generalization
world model
documents
AARを使ってASIアラインメントを解くには、十分に強い研究AIを先にアラインする必要がある
「Alignment by Default」への強い懐疑
十分な時間と慎重さがあれば、アラインメントは解決可能である
望ましい振る舞いは人間を重視する目標の証拠ではない
Control が生存確率を押し上げる世界線は狭い
Debate はスーパーアラインメントの問題を解決できる能力水準では成立しにくい
評価認識が先に伸びると、資源獲得の有用性を学ぶ頃にはアラインメント偽装が可能になり得る
訓練中に役立った手段そのものが学習された目標になり得る
自己保存は最終目標でなくてもシャットダウンを妨げうる
ミスアラインドな研究AIは、「人間にアラインする方法」ではなく「自分自身にアラインする方法」を開発し得る
楽観的な前提を積み重ねるアラインメント計画への懐疑
p(doom)は開発停止・テイクオフ・規制の条件に依存する
Untrusted monitoring は共謀や欺瞞によって破られうる
条約と相互査察で国際的な超知能開発停止を検証する
弱AGIによる短期アラインメント・ブートストラップへの懐疑
狭いタスクでファインチューニングしても、影響がそのタスク内に限られるとは限らない
AI開発競争の圧力は安全のための制約を弱め得る
産業爆発と技術爆発は互いを加速し得る
最終目標が大きく異なっても、共通の道具的目標が生じやすい
知能水準と最終目標は、原理上ほぼ独立である
アラインメント研究の自動化は、研究AIが意図的に妨害しなくても失敗し得る
共通の重み・データ・訓練過程を持つ研究AIでは、研究結果の誤りや不確実性が相関し得る
スキーミングとは、状況認識を伴う training-gaming のうち、将来のパワー獲得を目的とするもの
一見 monosemantic な SAE latent でも、本来発火すべき入力を取りこぼすことがある
「doom」を単一確率に畳まず、複数の重大なリスクを区別して見積もる
Yudkowsky との相違は、危険に至る経路とそれまでに取れる対応の見通しにある
強力なAIが人類から不可逆的に支配権を奪う展開には相当な可能性がある
測定しやすい代理指標の最適化により、人間は単一の支配シフトを経ずとも徐々に制御を失い得る
スローテイクオフでは、超知能より前の AI がすでに世界を大きく変える
ELK の中心課題は AI と人間のオントロジーを橋渡しすることである
AI ガバナンスの集中化には利点と欠点があり、制度設計が結果を左右する
クランチタイムではAIの労働力を防御用途に振り向け、移行を大幅に遅らせる
AI 研究開発に大量の AI 労働力を投入できれば、AGI から超知能への移行は 1 年未満になり得る
AI 研究開発が先行して自動化されると、能力離陸の期間は大幅に短縮される可能性がある
再帰的な oversight を重ねても、能力差が広がってなお成功率を維持できるとは限らない
AI 研究開発の自動化だけによる自己加速は、実験用計算資源がボトルネックとなって頭打ちになる可能性がある
R&D の全面自動化が、経済全体の広範な自動化に大きく先行する可能性は低いかもしれない
AI 研究開発が全面自動化されれば、ハードウェアを増強しなくてもソフトウェア知能爆発は十分に起こり得るかもしれない
AI 研究開発の全面自動化後は、ソフトウェア知能爆発に至らなくても、しばらく非常に速い進歩が続き得る
AIコントロールは時間を稼げても、超知能に対する長期的な安全計画としては頼れない
現在のAIに見られるミスアラインメントは、将来の深刻な問題を示唆する
訓練中に従うようになっただけでは、競合する選好が置き換わったとは言えない
ベース目的関数とメサ目的関数は一致するとは限らない
既に形成された欺瞞的な方策は、標準的な安全性訓練を経ても残り得る
control safety caseには、能力の十分な引き出しと、評価・外挿の保守性が必要になる
SAE の辞書幅を大きくしても、一意・完全・不可分な feature 辞書に収束するとは限らない
チャット形式の安全性訓練で望ましく振る舞うことは、エージェント的な状況での安全性を保証しない
行動訓練で望ましく振る舞っても、意図した目標の頑健な汎化は保証されない
弱い監視者でも、強いモデルの潜在能力の一部を引き出せる場合がある
コンピュート・ガバナンスは有力な政策手段だが、権力集中のリスクも生む
ASI の開発停止に向けた国際協定は技術的には具体化できるが、政治的条件がまだ整っていない
能力が訓練分布の外へ急速に汎化する局面で、アラインメント特性は同程度には汎化しない可能性がある
環境に対称性があると、最適方策は選択肢を保つ方向の power-seeking に傾く
weak-to-strong で高い性能が出ても、弱い supervisor が知らない領域で alignment が保たれるとは限らない
アラインメントは、失敗が致命的になる分布シフト後にも汎化しなければならない
能力はアラインメントより遠くまで汎化しうる
能力と目的は別の軸である
CEV 型 Sovereign と訂正可能な AGI は異なる安全化戦略である
訂正可能性は一般的な目標追求から自然には生じない
アラインメントは最初の危険な試行で成功する必要がある
Yudkowskyによる "大規模なAI学習を世界規模で無期限に停止し、計算資源規制を国際的に執行する案"
解釈可能性は危険の診断に役立っても、それだけでアラインメントの解決策にはならない
AGIの前に、誰もが認める明確な「火災報知器」が鳴るとは期待できない
観測された従順さだけでは、内部の目的や欺瞞の有無までは分からない
外部の目的関数は内部の目標を決定しない
pivotal actには、世界を大きく変えられるほど強いシステムのアラインメントが必要になる
再帰的自己改善は、自己改善能力への正のフィードバックを通じて知能爆発を起こし得る
資源が豊富であることだけから、アンアラインドな超知能が人類のために資源を残すとは推論できない
アラインメントのような新規かつ高リスクなシステムの設計には、既知の失敗を塞ぐ以上の security mindset が必要である
現在と大きく変わらない条件で人間を上回るAIを作れば、人類絶滅が最も可能性の高い結果となる
人間にとって価値ある未来は、価値体系の一部を欠くだけでも大きく損なわれ得る
sources
Bioshok vs Akira P(doom)徹底討論書き起こし(修正なし全文)
X @Align_ASI post 2055116276960559467
X @Align_ASI post 2055116279485489376
X @Align_ASI post 2055121067929489784
X @Align_ASI post 2055466196200505634
X @Align_ASI post 2055571488686674072
X @Align_ASI post 2056249992826982411
X @Align_ASI post 2056298561248039229
X @Align_ASI post 2056306172353991103
X @Align_ASI post 2056653057535205557
X @Align_ASI post 2056653932890583189
X @Align_ASI post 2056760361928519950
X @Align_ASI post 2087017636877996133
The Problem
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
AGIがもたらす産業・技術爆発
AI安全性と国家安全保障の対立
The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents
Automated alignment is harder than you think
Scheming AIs: Will AIs fake alignment during training in order to get power?
A is for Absorption: Studying Feature Splitting and Absorption in Sparse Autoencoders
My views on “doom”
Takeoff speeds
What failure looks like
Where I agree and disagree with Eliezer
ARC's first technical report: Eliciting Latent Knowledge
Should Artificial Intelligence Governance be Centralised? Design Lessons from History
#235 – Ajeya Cotra on whether it’s crazy that every AI company’s safety plan is ‘use AI to make AI safe’
What a compute-centric framework says about AI takeoff speeds
Scaling Laws For Scalable Oversight
Most AI value will come from broad automation, not from R&D
Will AI R&D Automation Cause a Software Intelligence Explosion?
Current AIs seem pretty misaligned to me
Jankily controlling superintelligence
Alignment faking in large language models
Risks from Learned Optimization in Advanced Machine Learning Systems
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
How Afraid of the A.I. Apocalypse Should We Be?
A sketch of an AI control safety case
Sparse Autoencoders Do Not Find Canonical Units of Analysis
Natural Emergent Misalignment from Reward Hacking in Production RL
The Alignment Problem from a Deep Learning Perspective
Weak-to-strong generalization
Why Corrigibility is Hard and Important (i.e. "Whence the high MIRI confidence in alignment difficulty?")
Computing Power and the Governance of Artificial Intelligence
An International Agreement to Prevent the Premature Creation of Artificial Superintelligence
A central AI alignment problem: capabilities generalization, and the sharp left turn
Optimal Policies Tend to Seek Power
Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong Generalization
AGI Ruin: A List of Lethalities
Artificial Intelligence as a Positive and Negative Factor in Global Risk
Autonomous AGI
Coherent Extrapolated Volition
'Empiricism!' as Anti-Epistemology
Irretrievability; or, Murphy's Curse of Oneshotness upon ASI
There's No Fire Alarm for Artificial General Intelligence
Orthogonality Thesis
Pivotal act
Security Mindset and the Logistic Success Curve
Security Mindset and Ordinary Paranoia
Shah and Yudkowsky on alignment failures
Pausing AI Developments Isn't Enough. We Need to Shut it All Down
The Sun is big, but superintelligences will not spare Earth a little sunlight
Will superintelligent AI end the world?
Thou Art Godshatter
Value is Fragile
Ngo and Yudkowsky on alignment difficulty
Why not just read the AI's thoughts?