Summary

Statistical inference in linear models centres on estimating relationships between a response variable and one or more predictors under the assumption that these relationships can be expressed as a linear combination. The ordinary least squares (OLS) framework provides point estimates of coefficients by minimising the sum of squared residuals, while classical inference relies on assumptions of homoscedasticity, normality of errors and independence. In practical settings—such as economics, genomics and social science—violations of these assumptions are common, motivating robust approaches. Developments over the past decade have extended inference to high-dimensional regimes, where the number of covariates grows with sample size, to semiparametric and nonparametric boundaries, and to models incorporating interactions and nonlinearities. Core objectives include obtaining valid confidence intervals, hypothesis tests for coefficient significance and predictive accuracy measures, even when error variances are unequal, covariates are many relative to observations, or interaction effects defy strict linearity. Computational advances and resampling techniques, such as bootstrap and jackknife, now complement analytical corrections, enabling practitioners to deliver reliable uncertainty quantification across diverse applications, from policy evaluation to risk modelling.

Research from Nature Portfolio

No recent Nature Portfolio content available.

Research from all publishers

Recent studies have introduced heteroscedasticity-robust covariance estimators tailored for linear regressions with a large number of control variables, demonstrating that traditional sandwich estimators can be inconsistent when covariate dimension scales with sample size, and proposing alternatives that restore correct test size and greater power. Parallel work has addressed the challenge of testing interaction effects in models that depart from linearity, offering curvilinear-robust procedures that maintain accurate false-positive rates and correct effect direction estimates even under complex response surfaces. In addition, advances in M-estimation software frameworks now allow researchers to specify general unbiased estimating equations and compute empirical sandwich variances, with extensions for finite-sample, heteroscedastic and autocorrelation adjustments, thus broadening the toolbox for inference beyond the classical OLS setting.

Statistical Inference in Linear Models publication trend

The graph below shows the total number of articles in statistical inference in linear models across all publications each year (not limited to Nature Index journals).

Technical terms

Linear Model: A statistical model in which the expected value of the dependent variable is expressed as a linear combination of explanatory variables and parameters.

Ordinary Least Squares (OLS): An estimation technique that obtains coefficient estimates by minimising the sum of squared differences between observed and predicted values.

Heteroscedasticity: A condition in which the variability of the error terms changes across levels of an explanatory variable or over observations.

Sandwich Estimator: A robust variance estimator for coefficient estimates that uses empirical residuals to adjust for model misspecification such as heteroscedasticity.

Interaction Term: A covariate constructed as the product of two or more predictors to model how the effect of one variable depends on the level of another.

References

  1. Interacting With Curves: How to Validly Test and Probe Interactions in the Real (Nonlinear) World. Advances in Methods and Practices in Psychological Science (2024).
  2. The Calculus of M-Estimation in R with geex.. Journal of Statistical Software (2020).
  3. Heteroscedasticity-Robust Inference in Linear Regression Models With Many Covariates. Journal of the American Statistical Association (2020).

About these summaries

This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.

Nature Strategy Reports
Turn complex research questions into confident strategic decisions 

When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.

  • Benchmark your performance against global peers using robust, methodologically sound analysis.

  • Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.

  • Gain tailored, decision-ready recommendations aligned to your strategic priorities.

Talk to us to learn more about our data dashboards and bespoke strategy reports.

Nature Masterclasses
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.

Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:

  • Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.

  • Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.

  • Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.

Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.