--- aliases: ja: - スケーラブルオーバーサイト concept_id: scalable_oversight preferred: en: scalable oversight ja: スケーラブル・オーバーサイト relations: relations: broader: - oversight prerequisite: - evaluation - oversight related: - weak_to_strong_generalization status: active --- # scalable oversight The problem and family of techniques for providing reliable supervision or training signals for AI systems when their outputs are too difficult for unaided human overseers to evaluate directly, especially as the systems become more capable than the overseers. ## broader - [oversight](../terminology/oversight.txt) ## prerequisite - [evaluation](../terminology/evaluation.txt) - [oversight](../terminology/oversight.txt) ## related - [weak-to-strong generalization](../terminology/weak_to_strong_generalization.txt) ## Documents: about - [再帰的な oversight を重ねても、能力差が広がってなお成功率を維持できるとは限らない](../documents/engels_et_al.recursive-oversight-does-not-automatically-scale-across-capability-gaps.txt) ## Documents: mentions - [Debate はスーパーアラインメントの問題を解決できる能力水準では成立しにくい](../documents/akira.debate-fails-at-superalignment-capability.txt) - [アラインメント研究の自動化は、研究AIが意図的に妨害しなくても失敗し得る](../documents/bowkis_et_al.automated-alignment-can-fail-without-deliberate-sabotage.txt) - [共通の重み・データ・訓練過程を持つ研究AIでは、研究結果の誤りや不確実性が相関し得る](../documents/bowkis_et_al.shared-training-can-correlate-automated-research-errors.txt) - [弱い監視者でも、強いモデルの潜在能力の一部を引き出せる場合がある](../documents/openai.weak-to-strong-supervision-can-recover-some-latent-capability.txt) - [weak-to-strong で高い性能が出ても、弱い supervisor が知らない領域で alignment が保たれるとは限らない](../documents/yang_et_al.weak-to-strong-success-can-hide-selective-misalignment.txt) [Corpus index](../index.txt)