Bayesian Model Selection in High-Dimensional Data

Summary

Bayesian model selection offers a coherent framework for identifying the most plausible models when the number of candidate predictors greatly exceeds the number of observations. Central to this approach is the comparison of models via their marginal likelihoods or Bayes factors, which naturally incorporate a penalty for over-complexity and embody Occam’s razor. In high-dimensional settings, priors that induce sparsity—such as spike-and-slab mixtures or continuous global-local shrinkage hierarchies—allow the data to inform which variables are truly important. Computational advances, including Markov chain Monte Carlo algorithms tailored for variable selection and fast variational inference schemes, have rendered fully Bayesian analyses practicable on modern datasets. Applications span genomics, where tens of thousands of gene expression measurements demand parsimonious models; neuroimaging, in which spatially correlated features require adaptive shrinkage; and finance, where large factor sets must be winnowed to forecast market movements. Recent work has emphasised scalable approximations to the marginal likelihood, novel prior constructions that balance sparsity with parameter interpretability, and diagnostic tools to assess convergence and robustness. Bayesian model selection thus unites statistical rigour with practical utility in the face of ever-growing data dimensionality.

Research from Nature Portfolio

Recent studies have demonstrated a computationally efficient method to evaluate Bayes factors in least-squares fitting, dispelling long-standing concerns over subjectivity and expense. The approach profiles the marginal likelihood via analytic approximations that quantify the trade-off between fit quality and model complexity, thereby operationalising Occam’s razor. It has been shown to discriminate effectively between models of equal parameter count and to discourage inclusion of spurious predictors. Practical implementations reveal improvements over classical information criteria, particularly in medium-sized datasets where physically meaningful parameters might otherwise be omitted or over-fitted.

Research from all publishers

An information-theoretic modification of the Bayesian Information Criterion has been proposed to control false discovery rates in genome-wide association studies. By incorporating an explicit penalty calibrated to the total number of genetic markers, the criterion balances statistical power against the risk of false positives, and heuristic search strategies enable its application to millions of variants. Comparative analyses indicate superior detection of truly associated loci with only a modest increase in type I error.

An adaptive ridge procedure offers a novel route to approximate L₀ penalisation for high-dimensional regression. Iteratively reweighted ridge regressions converge to sparse solutions that emulate exact subset selection, while avoiding the combinatorial complexity of direct L₀ optimisation. Extensive simulations demonstrate competitive performance against SCAD and adaptive LASSO in both orthogonal and non-orthogonal settings, with applications to segmentation of genomic data.

A theoretical study has elucidated how ranges of tuning parameters in various penalty functions influence the attainment of unbiased estimation and complete noise penalisation in ultra-high-dimensional contexts. By linking the minimum effect size and problem dimensionality to the choice of penalty family—especially those bridging L₀ and L₁ norms—this work guides practitioners in selecting regularisation schemes that maintain statistical guarantees while remaining computationally feasible.

Bayesian Model Selection in High-Dimensional Data publication trend

The graph below shows the total number of articles in bayesian model selection in high-dimensional data across all publications each year (not limited to Nature Index journals).

Technical terms

Bayes factor: Ratio of marginal likelihoods of two models, quantifying evidence in favour of one model over another.

High-dimensional data: Datasets in which the number of predictors substantially exceeds the number of observations, challenging classical inference.

Spike-and-slab prior: Mixture prior combining a point mass at zero (spike) with a continuous distribution (slab) to induce sparsity in parameter estimates.

Global-local shrinkage prior: Hierarchical prior structure that applies overall shrinkage while allowing individual parameters to escape heavy penalisation.

Variational inference: Deterministic approximation technique that converts Bayesian posterior estimation into an optimisation problem for scalable computation.

Marginal likelihood (model evidence): Integral of the likelihood over the prior, serving as the normalising constant and key component in Bayes factors.

References

  1. An Adaptive Ridge Procedure for L0 Regularization. PLOS ONE (2016).
  2. Analyzing Genome-Wide Association Studies with an FDR Controlling Modification of the Bayesian Information Criterion. PLOS ONE (2014).
  3. Easy computation of the Bayes factor to fully quantify Occam’s razor in least-squares fitting and to guide actions. Scientific Reports (2022).
  4. Designing penalty functions in high dimensional problems: The role of tuning parameters. Electronic Journal of Statistics (2016).

About these summaries

This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.

Nature Strategy Reports
Turn complex research questions into confident strategic decisions 

When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.

  • Benchmark your performance against global peers using robust, methodologically sound analysis.

  • Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.

  • Gain tailored, decision-ready recommendations aligned to your strategic priorities.

Talk to us to learn more about our data dashboards and bespoke strategy reports.

Nature Masterclasses
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.

Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:

  • Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.

  • Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.

  • Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.

Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.