Estimating the intrinsic dimension of datasets by a minimal neighborhood information

  • 2018-03-19 15:31:41
  • Elena Facco, Maria d'Errico, Alex Rodriguez, Alessandro Laio
  • 2

Abstract

Analyzing large volumes of high-dimensional data is an issue of fundamentalimportance in data science, molecular simulations and beyond. Severalapproaches work on the assumption that the important content of a datasetbelongs to a manifold whose Intrinsic Dimension (ID) is much lower than thecrude large number of coordinates. Such manifold is generally twisted andcurved, in addition points on it will be non-uniformly distributed: two factorsthat make the identification of the ID and its exploitation really hard. Herewe propose a new ID estimator using only the distance of the first and thesecond nearest neighbor of each point in the sample. This extreme minimalityenables us to reduce the effects of curvature, of density variation, and theresulting computational cost. The ID estimator is theoretically exact inuniformly distributed datasets, and provides consistent measures in general.When used in combination with block analysis, it allows discriminating therelevant dimensions as a function of the block size. This allows estimating theID even when the data lie on a manifold perturbed by a high-dimensional noise,a situation often encountered in real world data sets. We demonstrate theusefulness of the approach on molecular simulations and image analysis.

 

Quick Read (beta)

loading the full paper ...