Group Variable Selection in High-Dimensional Statistical Models
Summary
High-dimensional statistical models arise when the number of candidate predictors greatly exceeds the number of observations, a setting common in genomics, neuroimaging and finance. Group variable selection exploits known or hypothesised structure among predictors—such as biological pathways or functional modules—by applying penalties that act on predefined blocks of coefficients. Classic approaches include the group lasso, which enforces sparsity at the group level, and the sparse-group lasso, which combines group-wise selection with individual feature sparsity. Advances in theory have established conditions for consistency, oracle properties and prediction guarantees under both convex and nonconvex penalties. On the computational side, proximal algorithms, coordinate descent and alternating direction methods of multipliers have been tailored to handle very large data sets and complex penalty functions. By integrating prior information on group membership, these methods enhance interpretability and improve predictive accuracy. Applications span the identification of key gene clusters in cancer prognosis, the selection of relevant brain regions in neuroimaging studies and the extraction of meaningful factors in economic modelling. The global significance of this field lies in its ability to deliver parsimonious models that respect structured relationships among variables, thereby facilitating scientific discovery and translational insight.
Research from Nature Portfolio
Recent studies have extended meta-analytical frameworks for group variable selection by incorporating nonconvex penalties to synthesise evidence across multiple high-dimensional gene expression datasets. Three hierarchical decompositions of coefficients—based on half-norm, minimax concave and smoothly clipped absolute deviation penalties—offer flexible variable selection while preserving computational tractability. Efficient algorithms ensure convergence and scalability, and theoretical guarantees establish selection consistency. Application to heterogeneous lung cancer cohorts demonstrates enhanced reproducibility of biomarker discovery and clinically meaningful gene signatures, illustrating the power of nonconvex regularisation to integrate diverse studies and improve robustness.
Research from all publishers
New software implementations have enabled rapid deployment of sparse group regularisation in practical settings. A recent R package for the sparse group lasso offers highly optimised solution routines and warm-start strategies, allowing analysis of extremely large, sparse design matrices with both group- and within-group sparsity. Another development introduces knowledge distillation to sparse survival modelling: complex “teacher” survival models are used to train simpler “student” models under an ℓ1 penalty, reducing hyperparameter sensitivity and achieving competitive performance in high-dimensional time-to-event analyses through a user-friendly Python API. Parallel work on grouped logistic regression employs an lp,q regularisation scheme, establishing oracle inequalities under a group restricted eigenvalue condition. An alternating direction method of multipliers algorithm delivers efficient estimation, and numerical experiments in stock-factor selection and simulated data confirm accurate group-wise and individual sparsity recovery.
Group Variable Selection in High-Dimensional Statistical Models publication trend
The graph below shows the total number of articles in group variable selection in high-dimensional statistical models across all publications each year (not limited to Nature Index journals).
Technical terms
High-dimensional model: A regression or classification framework where the number of predictors exceeds or is comparable to the sample size, requiring special regularisation techniques.
Regularisation: The addition of a penalty term to a loss function to prevent overfitting and enforce sparsity or other structural constraints in estimated parameters.
Group lasso: A convex penalty that sums ℓ2 norms of coefficient blocks, encouraging entire groups of predictors to be selected or dropped together.
Sparse-group lasso: A hybrid penalty combining group-level ℓ2 norms with individual ℓ1 norms to achieve both block-wise and within-block sparsity.
Nonconvex penalty: A regularisation function, such as SCAD or MCP, that is not convex but can yield better bias–variance trade-offs and more accurate variable selection under certain conditions.
Knowledge distillation: A machine-learning technique in which a complex “teacher” model transfers predictive information to a simpler “student” model, facilitating regularised estimation and interpretability.
References
- sparsegl: An R Package for Estimating Sparse Group Lasso. Journal of Statistical Software (2024).
- sparsesurv: a Python package for fitting sparse survival models via knowledge distillation. Bioinformatics (2024).
- Meta-Analysis Based on Nonconvex Regularization. Scientific Reports (2020).
- Group Logistic Regression Models with lp,q Regularization. Mathematics (2022).
About these summaries
This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.
Turn complex research questions into confident strategic decisions
When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.
Benchmark your performance against global peers using robust, methodologically sound analysis.
Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.
Gain tailored, decision-ready recommendations aligned to your strategic priorities.
Talk to us to learn more about our data dashboards and bespoke strategy reports.
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.
Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:
Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.
Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.
Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.
Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.