B-test: A Non-parametric, Low Variance Kernel Two-sample Test

We propose a family of maximum mean discrepancy (MMD) kernel two-sample tests that have low sample complexity and are consistent. The test has a hyper-parameter that allows one to control the tradeoff between sample complexity and computational time. Our family of tests, which we denote as B-tests, is both computationally and statistically efficient, combining favorable properties of pre-viously proposed MMD two-sample tests. It does so by better leveraging sam-ples to produce low variance estimates in the finite sample case, while avoiding a quadratic number of kernel evaluations and complex null-hypothesis approxima-tion as would be required by tests relying on one sample U-statistics. The B-test uses a smaller than quadratic number of kernel evaluations and avoids completely the computational burden of complex null-hypothesis approximation, while main-taining consistency and probabilistically conservative thresholds on Type I error. Finally, recent results of combining multiple kernels transfer seamlessly to our hypothesis test, allowing a further increase in discriminative power and decrease in sample complexity. 1

Table for Estimating theGoodness of Fit of…Table for Estimating the Goodness of Fit of Empirical DistributionsApproximation Theoremsof Mathematical…Approximation Theorems of Mathematical Statistics.An Introduction to theBootstrapAn Introduction to the BootstrapOn a new multivariatetwo-sample testOn a new multivariate two-sample testReproducing KernelHilbert Spaces in…Reproducing Kernel Hilbert Spaces in Probability and StatisticsA Kernel Method for theTwo-Sample ProblemA Kernel Method for the Two-Sample ProblemTesting for Homogeneitywith Kernel Fisher…Testing for Homogeneity with Kernel Fisher Discriminant AnalysisA Kernel StatisticalTest of IndependenceA Kernel Statistical Test of IndependenceHilbert Space Embeddingsand Metrics on…Hilbert Space Embeddings and Metrics on Probability MeasuresOptimal kernel choicefor large-scale…Optimal kernel choice for large-scale two-sample testsA Kernel Two-Sample TestA Kernel Two-Sample TestA Scalable Bootstrap forMassive DataA Scalable Bootstrap for Massive DataM-Statistic for KernelChange-Point DetectionM-Statistic for Kernel Change-Point DetectionAdaptivity andComputation-Statistics…Adaptivity and Computation-Statistics Tradeoffs for Kernel and Distance based High Dimensional Two Sample TestingInterpretableDistribution Features…Interpretable Distribution Features with Maximum Testing PowerK2-ABC: ApproximateBayesian Computation…K2-ABC: Approximate Bayesian Computation with Kernel EmbeddingsLarge-scale kernelmethods for independenc…Large-scale kernel methods for independence testingAssessment of DataSuitability for Machine…Assessment of Data Suitability for Machine Prognosis Using Maximum Mean DiscrepancyDemystifying MMD GANsDemystifying MMD GANsDetecting and Correctingfor Label Shift with…Detecting and Correcting for Label Shift with Black Box PredictorsA Kernel MultipleChange-point Algorithm…A Kernel Multiple Change-point Algorithm via Model SelectionDistributional RandomForests: Heterogeneity…Distributional Random Forests: Heterogeneity Adjustment and Multivariate Distributional RegressionLearning Kernel TestsWithout Data SplittingLearning Kernel Tests Without Data SplittingAsymptotically OptimalOne- and Two-Sample…Asymptotically Optimal One- and Two-Sample Testing With KernelsB-test: ANon-parametric, Low…B-test: A Non-parametric, Low Variance Kernel Two-sample TestEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.