A significance test for the lasso

In the sparse linear regression setting, we consider testing the significance of the predictor variable that enters the current lasso model, in the sequence of models visited along the lasso solution path. We propose a simple test statistic based on lasso fitted values, called the covariance test statistic, and show that when the true model is linear, this statistic has an $\operatorname{Exp}(1)$ asymptotic distribution under the null hypothesis (the null being that all truly active variables are contained in the current lasso model). Our proof of this result for the special case of the first predictor to enter the model (i.e., testing for a single significant predictor variable against the global null) requires only weak assumptions on the predictor matrix $X$. On the other hand, our proof for a general step in the lasso path places further technical assumptions on $X$ and the generative model, but still allows for the important high-dimensional case $p>n$, and does not necessarily require that the current lasso model achieves perfect recovery of the truly active variables. Of course, for testing the significance of an additional variable between two nested linear models, one typically uses the chi-squared test, comparing the drop in residual sum of squares (RSS) to a $\chi^{2}_{1}$ distribution. But when this additional variable is not fixed, and has been chosen adaptively or greedily, this test is no longer appropriate: adaptivity makes the drop in RSS stochastically much larger than $\chi^{2}_{1}$ under the null hypothesis. Our analysis explicitly accounts for adaptivity, as it must, since the lasso builds an adaptive sequence of linear models as the tuning parameter $\lambda$ decreases. In this analysis, shrinkage plays a key role: though additional variables are chosen adaptively, the coefficients of lasso active variables are shrunken due to the $\ell_{1}$ penalty. Therefore, the test statistic (which is based on lasso fitted values) is in a sense balanced by these two opposing properties—adaptivity and shrinkage—and its null distribution is tractable and asymptotically $\operatorname{Exp}(1)$.

L1-Regularization PathAlgorithm for…L1-Regularization Path Algorithm for Generalized Linear ModelsHIGH DIMENSIONALVARIABLE SELECTIONHIGH DIMENSIONAL VARIABLE SELECTIONp-Values forHigh-Dimensional…p-Values for High-Dimensional RegressionNESTA: A Fast andAccurate First-Order…NESTA: A Fast and Accurate First-Order Method for Sparse RecoveryRegularization Paths forGeneralized Linear…Regularization Paths for Generalized Linear Models via Coordinate Descent.Stability SelectionStability SelectionA Perturbation Methodfor Inference on…A Perturbation Method for Inference on Regularized Regression EstimatesScaled sparse linearregressionScaled sparse linear regressionDegrees of freedom inlasso problemsDegrees of freedom in lasso problemsConfidence Intervals forLow Dimensional…Confidence Intervals for Low Dimensional Parameters in High Dimensional Linear ModelsThe lasso problem anduniquenessThe lasso problem and uniquenessConfidence Intervals andHypothesis Testing for…Confidence Intervals and Hypothesis Testing for High-Dimensional RegressionConfidence Intervals andHypothesis Testing for…Confidence Intervals and Hypothesis Testing for High-Dimensional Statistical ModelsConfidence Intervals andHypothesis Testing for…Confidence Intervals and Hypothesis Testing for High-Dimensional RegressionOptimal Inference AfterModel SelectionOptimal Inference After Model SelectionHigh-DimensionalInference: Confidence…High-Dimensional Inference: Confidence Intervals, p-Values and R-Software hdiUsing Lasso forPredictor Selection and…Using Lasso for Predictor Selection and to Assuage Overfitting: A Method Long Overlooked in Behavioral SciencesSLOPE—Adaptive variableselection via convex…SLOPE—Adaptive variable selection via convex optimizationExact Post-SelectionInference for Sequentia…Exact Post-Selection Inference for Sequential Regression ProceduresA general theory ofhypothesis tests and…A general theory of hypothesis tests and confidence regions for sparse high dimensional modelsThe Induced Smoothedlasso: A practical…The Induced Smoothed lasso: A practical framework for hypothesis testing in high dimensional regressionBootstrapping and samplesplitting for…Bootstrapping and sample splitting for high-dimensional, assumption-lean inferenceMarkov NeighborhoodRegression for…Markov Neighborhood Regression for High-Dimensional InferenceStatistical Significancein High-dimensional…Statistical Significance in High-dimensional Linear Mixed ModelsA significance test forthe lassoA significance test for the lassoEarlier referencesFocus paperCiting papersOlderNewer

Click a node to pin it, click the empty canvas to go back to this paper, or hover to preview. Open a node’s page from its title.