Energy Transformer

Our work combines aspects of three promising paradigms in machine learning, namely, attention mechanism, energy-based models, and associative memory. Attention is the power-house driving modern deep learning successes, but it lacks clear theoretical foundations. Energy-based models allow a principled approach to discriminative and generative tasks, but the design of the energy functional is not straightforward. At the same time, Dense Associative Memory models or Modern Hopfield Networks have a well-established theoretical foundation, and allow an intuitive design of the energy function. We propose a novel architecture, called the Energy Transformer (or ET for short), that uses a sequence of attention layers that are purposely designed to minimize a specifically engineered energy function, which is responsible for representing the relationships between the tokens. In this work, we introduce the theoretical foundations of ET, explore its empirical capabilities using the image completion task, and obtain strong quantitative results on the graph anomaly detection and graph classification tasks.

Semi-SupervisedClassification with…Semi-Supervised Classification with Graph Convolutional NetworksSGDR: StochasticGradient Descent with…SGDR: Stochastic Gradient Descent with Warm RestartsGraph Attention NetworksGraph Attention NetworksASAP: Adaptive StructureAware Pooling for…ASAP: Adaptive Structure Aware Pooling for Learning Hierarchical Graph RepresentationsA Generalization ofTransformer Networks to…A Generalization of Transformer Networks to GraphsHierarchical AssociativeMemoryHierarchical Associative MemoryHopfield Networks is AllYou NeedHopfield Networks is All You NeedLarge Associative MemoryProblem in Neurobiology…Large Associative Memory Problem in Neurobiology and Machine LearningRethinking Graph NeuralNetworks for Anomaly…Rethinking Graph Neural Networks for Anomaly DetectionMetaFormer is ActuallyWhat You Need for VisionMetaFormer is Actually What You Need for VisionA Survey of TransformersA Survey of TransformersTransformers from anOptimization PerspectiveTransformers from an Optimization PerspectiveOn Sparse ModernHopfield ModelOn Sparse Modern Hopfield ModelEnergy-Based CrossAttention for Bayesian…Energy-Based Cross Attention for Bayesian Context Update in Text-to-Image Diffusion ModelsEnd-to-endDifferentiable…End-to-end Differentiable Clustering with Associative MemoriesBiSHop: Bi-DirectionalCellular Learning for…BiSHop: Bi-Directional Cellular Learning for Tabular Data with Generalized Sparse Modern Hopfield ModelOutlier-EfficientHopfield Layers for…Outlier-Efficient Hopfield Layers for Large Transformer-Based ModelsSTanHop: Sparse TandemHopfield Model for…STanHop: Sparse Tandem Hopfield Model for Memory-Enhanced Time Series PredictionOn Computational Limitsof Modern Hopfield…On Computational Limits of Modern Hopfield Models: A Fine-Grained Complexity AnalysisProvably Optimal MemoryCapacity for Modern…Provably Optimal Memory Capacity for Modern Hopfield Models: Transformer-Compatible Dense Associative Memories as Spherical CodesConv-CoA: ImprovingOpen-domain Question…Conv-CoA: Improving Open-domain Question Answering in Large Language Models via Conversational Chain-of-ActionTree Attention:Topology-aware Decoding…Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clustersMachine learning inbiological physics: Fro…Machine learning in biological physics: From biomolecular prediction to designNonparametric ModernHopfield ModelsNonparametric Modern Hopfield ModelsEnergy TransformerEnergy Transformer過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。