The Platonic Representation Hypothesis

We argue that representations in AI models, particularly deep networks, are converging. First, we survey many examples of convergence in the literature: over time and across multiple domains, the ways by which different neural networks represent data are becoming more aligned. Next, we demonstrate convergence across data modalities: as vision models and language models get larger, they measure distance between datapoints in a more and more alike way. We hypothesize that this convergence is driving toward a shared statistical model of reality, akin to Plato's concept of an ideal reality. We term such a representation the platonic representation and discuss several possible selective pressures toward it. Finally, we discuss the implications of these trends, their limitations, and counterexamples to our analysis.

Representation Learningwith Contrastive…Representation Learning with Contrastive Predictive CodingBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingAn Image is Worth 16x16Words: Transformers for…An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleOn the Opportunities andRisks of Foundation…On the Opportunities and Risks of Foundation ModelsLlama 2: Open Foundationand Fine-Tuned Chat…Llama 2: Open Foundation and Fine-Tuned Chat ModelsGemini: A Family ofHighly Capable…Gemini: A Family of Highly Capable Multimodal ModelsPaLM-E: An EmbodiedMultimodal Language…PaLM-E: An Embodied Multimodal Language ModelMistral 7BMistral 7BGit Re-Basin: MergingModels modulo…Git Re-Basin: Merging Models modulo Permutation SymmetriesGPT-4 Technical ReportGPT-4 Technical ReportHelping Cancer Patientsto Choose the Best…Helping Cancer Patients to Choose the Best Treatment: Towards Automated Data-Driven and Personalized Information Presentation of Cancer Treatment OptionsMixtral of ExpertsMixtral of ExpertsGetting aligned onrepresentational…Getting aligned on representational alignmentGeneralization fromStarvation: Hints of…Generalization from Starvation: Hints of Universality in LLM Knowledge Graph LearningMATE: Meet At TheEmbedding - Connecting…MATE: Meet At The Embedding - Connecting Images with Long TextsI Learn Better If YouSpeak My Language…I Learn Better If You Speak My Language: Understanding the Superior Performance of Fine-Tuning Large Language Models with LLM-Generated ResponsesNeural Scaling LawsRooted in the Data…Neural Scaling Laws Rooted in the Data DistributionAstroPT: Scaling LargeObservation Models for…AstroPT: Scaling Large Observation Models for AstronomyWe're Different, We'rethe Same: Creative…We're Different, We're the Same: Creative Homogeneity Across LLMsLayers at Similar DepthsGenerate Similar…Layers at Similar Depths Generate Similar Activations Across LLM ArchitecturesAutomating the Searchfor Artificial Life wit…Automating the Search for Artificial Life with Foundation ModelsActivation SpaceInterventions Can Be…Activation Space Interventions Can Be Transferred Between Large Language ModelsUncovering theComputational…Uncovering the Computational Ingredients of Human-Like Representations in LLMsDisentangling theFactors of Convergence…Disentangling the Factors of Convergence between Brains and Computer Vision ModelsThe PlatonicRepresentation…The Platonic Representation Hypothesis過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。