Toy Models of Superposition

Demonstrates that hallucination in neural networks is a structural consequence of ungoverned superposition. Introduces a governance filter operating on the Gram matrix of the weight matrix that eliminates 100% of false-positive activations across all tested sparsity regimes (a toy-model feasibility result; true-positive retention not yet reported) in the Elhage et al. (2022) toy model framework, at a cost of 4-31% increased reconstruction error. Proposes a hierarchical two-supervisor architecture (routing supervisor + governance supervisor). No retraining, no extra parameters, applied as a post-hoc architectural layer. Originally a Google contest submission. --- **Author:** James E. Dunn — Independent Researcher, Hydrogen Lifecycle Research Programme **ORCID:** https://orcid.org/0009-0005-2679-6574 **Corpus (author search):** https://zenodo.org/search?q=creators.orcid:0009-0005-2679-6574 **SciX (NASA discovery):** https://scixplorer.org/search/q=orcid%3A0009-0005-2679-6574 **ADS (Harvard-CfA):** https://ui.adsabs.harvard.edu/search/q=orcid%3A0009-0005-2679-6574 **License:** CC BY 4.0 International

Visualizing andUnderstanding Recurrent…Visualizing and Understanding Recurrent NetworksObject Detectors Emergein Deep Scene CNNsObject Detectors Emerge in Deep Scene CNNsUnsupervisedRepresentation Learning…Unsupervised Representation Learning with Deep Convolutional Generative Adversarial NetworksTowards the Science ofSecurity and Privacy in…Towards the Science of Security and Privacy in Machine LearningLearning to GenerateReviews and Discovering…Learning to Generate Reviews and Discovering SentimentDelving intoTransferable Adversaria…Delving into Transferable Adversarial Examples and Black-box AttacksOn the importance ofsingle directions for…On the importance of single directions for generalizationAdversarial SpheresAdversarial SpheresLearningPerceptually-Aligned…Learning Perceptually-Aligned Representations via Adversarial RobustnessGrokking: GeneralizationBeyond Overfitting on…Grokking: Generalization Beyond Overfitting on Small Algorithmic DatasetsA Review of SparseExpert Models in Deep…A Review of Sparse Expert Models in Deep LearningToward Transparent AI: ASurvey on Interpreting…Toward Transparent AI: A Survey on Interpreting the Inner Structures of Deep Neural NetworksLanguage ModelsRepresent Space and TimeLanguage Models Represent Space and TimeThe Missing CurveDetectors of…The Missing Curve Detectors of InceptionV1: Applying Sparse Autoencoders to InceptionV1 Early VisionDisentangling DenseEmbeddings with Sparse…Disentangling Dense Embeddings with Sparse AutoencodersScaling and evaluatingsparse autoencodersScaling and evaluating sparse autoencodersA is for Absorption:Studying Feature…A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersLearning Multi-LevelFeatures with Matryoshk…Learning Multi-Level Features with Matryoshka Sparse AutoencodersUniversal SparseAutoencoders…Universal Sparse Autoencoders: Interpretable Cross-Model Concept AlignmentAutomaticallyInterpreting Millions o…Automatically Interpreting Millions of Features in Large Language ModelsPersona Features ControlEmergent MisalignmentPersona Features Control Emergent MisalignmentInterpreting visiontransformers via…Interpreting vision transformers via residual replacement modelModel Editing as aRobust and Denoised…Model Editing as a Robust and Denoised variant of DPO: A Case Study on ToxicityToy Models ofSuperpositionToy Models of Superposition過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。