Relative representations enable zero-shot latent space communication

Neural networks embed the geometric structure of a data manifold lying in a high-dimensional space into latent representations. Ideally, the distribution of the data points in the latent space should depend only on the task, the data, the loss, and other architecture-specific constraints. However, factors such as the random weights initialization, training hyperparameters, or other sources of randomness in the training phase may induce incoherent latent spaces that hinder any form of reuse. Nevertheless, we empirically observe that, under the same data and modeling choices, the angles between the encodings within distinct latent spaces do not change. In this work, we propose the latent similarity between each sample and a fixed set of anchors as an alternative data representation, demonstrating that it can enforce the desired invariances without any additional training. We show how neural architectures can leverage these relative representations to guarantee, in practice, invariance to latent isometries and rescalings, effectively enabling latent space communication: from zero-shot model stitching to latent space comparison between diverse settings. We extensively validate the generalization capability of our approach on different datasets, spanning various modalities (images, text, graphs), tasks (e.g., classification, reconstruction) and architectures (e.g., CNNs, GCNs, transformers).

CollectiveClassification in…Collective Classification in Network DataFashion-MNIST: a NovelImage Dataset for…Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning AlgorithmsBERT: Pre-training ofDeep Bidirectional…BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingTransformers:State-of-the-Art Natura…Transformers: State-of-the-Art Natural Language ProcessingThe Multilingual AmazonReviews CorpusThe Multilingual Amazon Reviews CorpusAre All Good Word VectorSpaces Isomorphic?Are All Good Word Vector Spaces Isomorphic?Fantastic Embeddings andHow to Align Them…Fantastic Embeddings and How to Align Them: Zero-Shot Inference in a Multi-Shop ScenarioSelf-Attention BetweenDatapoints: Going Beyon…Self-Attention Between Datapoints: Going Beyond Individual Input-Output Pairs in Deep LearningWikiMatrix: Mining 135MParallel Sentences in…WikiMatrix: Mining 135M Parallel Sentences in 1620 Language Pairs from WikipediaTheSelf-Optimal-Transport…The Self-Optimal-Transport Feature TransformASIF: Coupled Data TurnsUnimodal Models to…ASIF: Coupled Data Turns Unimodal Models to Multimodal Without TrainingCoReS: CompatibleRepresentations via…CoReS: Compatible Representations via StationarityInference-TimeIntervention: Eliciting…Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelAlignment with humanrepresentations support…Alignment with human representations supports robust few-shot learningOn the Origins of LinearRepresentations in Larg…On the Origins of Linear Representations in Large Language ModelsA Survey onHallucination in Large…A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open QuestionsThe PlatonicRepresentation…The Platonic Representation HypothesisLearning InterpretableConcepts: Unifying…Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation ModelsFostering effectivehybrid human-LLM…Fostering effective hybrid human-LLM reasoning and decision makingHarnessing the UniversalGeometry of EmbeddingsHarnessing the Universal Geometry of EmbeddingsThe Geometry ofCategorical and…The Geometry of Categorical and Hierarchical Concepts in Large Language ModelsCross-Entropy Is All YouNeed To Invert the Data…Cross-Entropy Is All You Need To Invert the Data Generating ProcessI Predict Therefore IAm: Is Next Token…I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?There Will Be aScientific Theory of…There Will Be a Scientific Theory of Deep LearningRelative representationsenable zero-shot latent…Relative representations enable zero-shot latent space communication過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。