Large-scale kernel methods for independence testing

Representations of probability measures in reproducing kernel Hilbert spaces provide a flexible framework for fully nonparametric hypothesis tests of independence, which can capture any type of departure from independence, including nonlinear associations and multivariate interactions. However, these approaches come with an at least quadratic computational cost in the number of observations, which can be prohibitive in many applications. Arguably, it is exactly in such large-scale datasets that capturing any type of dependence is of interest, so striking a favourable tradeoff between computational efficiency and test performance for kernel independence tests would have a direct impact on their applicability in practice. In this contribution, we provide an extensive study of the use of large-scale kernel approximations in the context of independence testing, contrasting block-based, Nystrom and random Fourier feature approaches. Through a variety of synthetic data experiments, it is demonstrated that our novel large scale methods give comparable performance with existing methods whilst using significantly less computation time and memory.

A Kernel StatisticalTest of IndependenceA Kernel Statistical Test of IndependenceKernel Measures ofConditional DependenceKernel Measures of Conditional DependenceA Fast, ConsistentKernel Two-Sample TestA Fast, Consistent Kernel Two-Sample TestKernel-based ConditionalIndependence Test and…Kernel-based Conditional Independence Test and Application in Causal DiscoveryEquivalence ofdistance-based and…Equivalence of distance-based and RKHS-based statistics in hypothesis testingA Kernel Two-Sample TestA Kernel Two-Sample TestOptimal kernel choicefor large-scale…Optimal kernel choice for large-scale two-sample testsA Kernel Test forThree-Variable…A Kernel Test for Three-Variable InteractionsThe RandomizedDependence CoefficientThe Randomized Dependence CoefficientDistance covariance inmetric spacesDistance covariance in metric spacesFastMMD: Ensemble ofCircular Discrepancy fo…FastMMD: Ensemble of Circular Discrepancy for Efficient Two-Sample TestGaussian Processes forIndependence Tests with…Gaussian Processes for Independence Tests with Non-iid Data in Causal Inferenceopenalex_id:w2963535485openalex_id:w2963535485The Chi-Square Test ofDistance CorrelationThe Chi-Square Test of Distance CorrelationGaussian Processes andKernel Methods: A Revie…Gaussian Processes and Kernel Methods: A Review on Connections and EquivalencesThe Exact Equivalence ofDistance and Kernel…The Exact Equivalence of Distance and Kernel Methods for Hypothesis TestingApproximate Kernel-BasedConditional Independenc…Approximate Kernel-Based Conditional Independence Tests for Fast Non-Parametric Causal DiscoveryA fast algorithm forcomputing distance…A fast algorithm for computing distance correlationDiscovering anddeciphering…Discovering and deciphering relationships across disparate data modalitiesMeasuring Association onTopological Spaces Usin…Measuring Association on Topological Spaces Using Kernels and Geometric GraphsMore Powerful SelectiveKernel Tests for Featur…More Powerful Selective Kernel Tests for Feature SelectionLearning withHilbert-Schmidt…Learning with Hilbert-Schmidt independence criterion: A review and new perspectivesChatterjee CorrelationCoefficient: A robust…Chatterjee Correlation Coefficient: A robust alternative for classic correlation methods in geochemical studies- (including “TripleCpy” Python package)A survey of some recentdevelopments in measure…A survey of some recent developments in measures of associationLarge-scale kernelmethods for independenc…Large-scale kernel methods for independence testingEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.