--- aliases: en: - AI interpretability - ML interpretability concept_id: interpretability preferred: en: interpretability ja: 解釈可能性 status: active --- # interpretability The study and use of methods for making the behavior, decision processes, representations, or internal computations of AI and machine-learning systems understandable to human or other external observers. ## Documents: about - [解釈可能性は危険の診断に役立っても、それだけでアラインメントの解決策にはならない](../documents/yudkowsky.interpretability-is-diagnostic-not-by-itself-an-alignment-solution.txt) ## Documents: mentions - [一見 monosemantic な SAE latent でも、本来発火すべき入力を取りこぼすことがある](../documents/chanin_et_al.monosemantic-sae-features-can-have-systematic-false-negatives.txt) - [SAE の辞書幅を大きくしても、一意・完全・不可分な feature 辞書に収束するとは限らない](../documents/leask_et_al.sae-dictionaries-need-not-be-canonical.txt) [Corpus index](../index.txt)