Variable Selection and Subgroup Analysis in High-Dimensional Data

Summary

Modern scientific investigations frequently generate datasets with thousands to millions of features, ranging from genomic profiles to sensor measurements. Analysing such high-dimensional data demands two complementary tasks. Variable selection isolates the most informative predictors by imposing sparsity through penalisation or shrinkage techniques, thereby improving interpretability and predictive accuracy. Common approaches include Lasso, elastic net and group-Lasso penalties, which automatically eliminate irrelevant variables and guard against overfitting. Subgroup analysis then seeks to uncover latent clusters or patient groups that exhibit distinct patterns of association between covariates and outcomes. This can be achieved via convex clustering, regularised k-means or model-based mixture models, often augmented by penalties that encourage both sparsity and coherent grouping. Together, these methods provide a principled framework for discovering biomarkers, defining molecular subtypes in cancer or tailoring interventions in precision medicine. Recent advances have focused on integrating prior information, leveraging multi-omics measurements and ensuring theoretical guarantees such as selection consistency and accurate subgroup recovery. Scalable optimisation algorithms—including variants of the alternating direction method of multipliers—enable practical application to ever-larger datasets in biology, economics and engineering, emphasising global relevance and real-world impact.

Research from Nature Portfolio

No recent Nature Portfolio content available.

Research from all publishers

Advances in sparse convex clustering have enabled more accurate disease subtyping by integrating external knowledge. One method incorporates text-mined literature through group-Lasso penalties to guide feature selection and cluster formation, allowing the joint analysis of multi-omics datasets and yielding improved subtype delineation in breast and lung cancer. Seminal work on regularised k-means extended classical clustering to high-dimensional settings by adding an adaptive group-Lasso penalty on cluster centres, demonstrating both variable elimination and asymptotic consistency as the number of features grows. More recently, hierarchical clustering frameworks have been devised to handle heterogeneous feature types and prior pairwise relationships, first performing a rough split and then refining clusters to reveal nested subgroup structures. These algorithms employ concave fusion penalties and ensure statistical consistency, offering deeper insights into complex biological processes such as tumour heterogeneity by flexibly integrating diverse data sources.

Variable Selection and Subgroup Analysis in High-Dimensional Data publication trend

The graph below shows the total number of articles in variable selection and subgroup analysis in high-dimensional data across all publications each year (not limited to Nature Index journals).

Technical terms

High-dimensional data: Datasets with a number of features that is large relative to the sample size, often leading to challenges in estimation and interpretation.

Variable selection: The process of identifying a subset of relevant predictors by imposing sparsity, commonly via penalised regression techniques.

Subgroup analysis: The discovery of latent clusters or subpopulations within data that exhibit distinct relationships among variables or outcomes.

Regularisation: A strategy that adds penalty terms to an objective function to control model complexity and prevent overfitting.

Convex clustering: A clustering approach framed as a convex optimisation problem, which promotes cluster fusion via fusion or total-variation penalties.

References

  1. Information-incorporated sparse convex clustering for disease subtyping. Bioinformatics (2023).
  2. Clustering on hierarchical heterogeneous data with prior pairwise relationships. BMC Bioinformatics (2024).
  3. Regularized k-means clustering of high-dimensional data and its asymptotic consistency. Electronic Journal of Statistics (2012).
  4. A parallel ADMM-based convex clustering method. EURASIP Journal on Advances in Signal Processing (2022).

About these summaries

This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.

Nature Strategy Reports
Turn complex research questions into confident strategic decisions 

When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.

  • Benchmark your performance against global peers using robust, methodologically sound analysis.

  • Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.

  • Gain tailored, decision-ready recommendations aligned to your strategic priorities.

Talk to us to learn more about our data dashboards and bespoke strategy reports.

Nature Masterclasses
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.

Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:

  • Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.

  • Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.

  • Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.

Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.