Intrinsic dimension of data representations in deep neural networks

Deep neural networks progressively transform their inputs across multiple processing layers. What are the geometrical properties of the representations learned by these networks? Here we study the intrinsic dimensionality (ID) of data representations, i.e. the minimal number of parameters needed to describe a representation. We find that, in a trained network, the ID is orders of magnitude smaller than the number of units in each layer. Across layers, the ID first increases and then progressively decreases in the final layers. Remarkably, the ID of the last hidden layer predicts classification accuracy on the test set. These results can neither be found by linear dimensionality estimates (e.g., with principal component analysis), nor in representations that had been artificially linearized. They are neither found in untrained networks, nor in networks that are trained on randomized labels. This suggests that neural networks that can generalize are those that transform the data into low-dimensional, but not necessarily flat manifolds.

Optimal Brain DamageOptimal Brain DamageMaximum LikelihoodEstimation of Intrinsic…Maximum Likelihood Estimation of Intrinsic DimensionSparse coding of sensoryinputsSparse coding of sensory inputsUntangling invariantobject recognitionUntangling invariant object recognitionImageNet: A large-scalehierarchical image…ImageNet: A large-scale hierarchical image databaseImageNet Classificationwith Deep Convolutional…ImageNet Classification with Deep Convolutional Neural NetworksPredicting Parameters inDeep LearningPredicting Parameters in Deep LearningVery Deep ConvolutionalNetworks for Large-Scal…Very Deep Convolutional Networks for Large-Scale Image RecognitionEstimating the intrinsicdimension of datasets b…Estimating the intrinsic dimension of datasets by a minimal neighborhood informationClassification andGeometry of General…Classification and Geometry of General Perceptual ManifoldsAutomaticdifferentiation in…Automatic differentiation in PyTorchNonlinear processing ofshape information in ra…Nonlinear processing of shape information in rat lateral extrastriate cortexIntrinsic dimensionestimation for locally…Intrinsic dimension estimation for locally undersampled dataA Neural Scaling Lawfrom the Dimension of…A Neural Scaling Law from the Dimension of the Data ManifoldStatistical learningtheory of structured…Statistical learning theory of structured dataTriple descent and thetwo kinds of…Triple descent and the two kinds of overfitting: Where & why do they appear?Relative stabilitytoward diffeomorphisms…Relative stability toward diffeomorphisms in deep nets indicates performanceSolvable Model for theLinear Separability of…Solvable Model for the Linear Separability of Structured DataHigh-performing neuralnetwork models of visua…High-performing neural network models of visual cortex benefit from high latent dimensionalityDADApy: Distance-basedAnalysis of…DADApy: Distance-based Analysis of DAta-manifolds in PythonDeep neural networksarchitectures from the…Deep neural networks architectures from the perspective of manifold learningStatistical Mechanics:Theory and Molecular…Statistical Mechanics: Theory and Molecular SimulationIntrinsic DimensionEstimation for Robust…Intrinsic Dimension Estimation for Robust Detection of AI-Generated TextsRepresentations andgeneralization in…Representations and generalization in artificial and brain neural networksIntrinsic dimension ofdata representations in…Intrinsic dimension of data representations in deep neural networks過去の参考文献中心の論文この論文を引用する論文古い新しい

ノードをクリックするとフォーカスを固定、空白をクリックすると本論文に戻ります。ホバーで一時的にプレビューできます。各ノードのページはタイトルから開けます。