A Kernel Two-Sample Test

We propose a framework for analyzing and comparing distributions, which we use to construct statistical tests to determine if two samples are drawn from different distributions. Our test statistic is the largest difference in expectations over functions in the unit ball of a reproducing kernel Hilbert space (RKHS), and is called the maximum mean discrepancy (MMD).We present two distributionfree tests based on large deviation bounds for the MMD, and a third test based on the asymptotic distribution of this statistic. The MMD can be computed in quadratic time, although efficient linear time approximations are available. Our statistic is an instance of an integral probability metric, and various classical metrics on distributions are obtained when alternative function classes are used in place of an RKHS. We apply our two-sample tests to a variety of problems, including attribute matching for databases using the Hungarian marriage method, where they perform strongly. Excellent performance is also obtained when comparing distributions over graphs, for which these are the first such tests.

A Kernel Method for theTwo-Sample ProblemA Kernel Method for the Two-Sample ProblemIntegrating structuredbiological data by…Integrating structured biological data by Kernel Maximum Mean DiscrepancyA Hilbert SpaceEmbedding for…A Hilbert Space Embedding for DistributionsKernel Measures ofConditional DependenceKernel Measures of Conditional DependenceA Kernel StatisticalTest of IndependenceA Kernel Statistical Test of IndependenceA Kernel Approach toComparing DistributionsA Kernel Approach to Comparing DistributionsInjective Hilbert SpaceEmbeddings of…Injective Hilbert Space Embeddings of Probability MeasuresA Fast, ConsistentKernel Two-Sample TestA Fast, Consistent Kernel Two-Sample TestKernel Choice andClassifiability for RKH…Kernel Choice and Classifiability for RKHS Embeddings of Probability DistributionsHilbert Space Embeddingsand Metrics on…Hilbert Space Embeddings and Metrics on Probability MeasuresUniversality,Characteristic Kernels…Universality, Characteristic Kernels and RKHS Embedding of MeasuresLearning in Hilbert vs.Banach Spaces: A Measur…Learning in Hilbert vs. Banach Spaces: A Measure Embedding ViewpointDeep Learning ofTransferable…Deep Learning of Transferable Representation for Scalable Domain AdaptationRevisiting BatchNormalization For…Revisiting Batch Normalization For Practical Domain AdaptationNonparametric Detectionof Geometric Structures…Nonparametric Detection of Geometric Structures Over NetworksDomain Generalizationand Adaptation Using Lo…Domain Generalization and Adaptation Using Low Rank Exemplar SVMsKernel DistributionEmbeddings: Universal…Kernel Distribution Embeddings: Universal Kernels, Characteristic Kernels and Kernel Metrics on DistributionsA SequentialNon-Parametric…A Sequential Non-Parametric Multivariate Two-Sample TestSinkhorn AutoEncodersSinkhorn AutoEncodersManifold CriterionGuided Transfer Learnin…Manifold Criterion Guided Transfer Learning via Intermediate Domain GenerationImproved TrAdaBoost andits Application to…Improved TrAdaBoost and its Application to Transaction Fraud DetectionRethinking ImportanceWeighting for Deep…Rethinking Importance Weighting for Deep Learning under Distribution ShiftEEG-Based EmotionRecognition Using…EEG-Based Emotion Recognition Using Regularized Graph Neural NetworksGeneralizing to UnseenDomains: A Survey on…Generalizing to Unseen Domains: A Survey on Domain GeneralizationA Kernel Two-Sample TestA Kernel Two-Sample TestEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.