Variable Selection and Estimation in Survival Analysis
Summary
Survival analysis examines time-to-event data, addressing situations in which the outcome is the time until an event of interest, such as death, disease recurrence or machine failure. A central challenge lies in managing incomplete observations, or censoring, and in selecting the most informative covariates when many predictors are available. Variable selection reduces model complexity and enhances interpretability, while estimation focuses on obtaining unbiased and efficient estimates of covariate effects. Classical approaches, such as stepwise selection in the Cox proportional hazards model, often falter in high-dimensional settings. Over the last decade, regularisation techniques—including Lasso, elastic net and group penalties—have become standard tools to impose sparsity and to guard against overfitting. Extensions of the Cox model, such as additive hazards and accelerated failure time models, offer alternative hazard structures that may suit particular applications. Recent methodological advances integrate bias correction for left truncation, account for interval censoring or length-biased sampling, and allow simultaneous group and within‐group selection. Together, these developments have broadened the applicability of survival models to large-scale genomic studies, electronic health records and complex observational cohorts, delivering robust risk prediction, feature ranking and insight into disease mechanisms.
Research from Nature Portfolio
No recent Nature Portfolio content available.
Research from all publishers
Penalised regression methods have been adapted to address left-truncated and right-censored survival data. A recent study applied a penalised Cox proportional hazards model that incorporates left-truncation adjustment, demonstrating through simulations and a real-world clinico-genomic database that failure to correct for entry bias can substantially distort coefficient estimates. This work underscores the importance of combining regularisation with bias correction when dealing with delayed study entry.
Advances in additive hazards modelling have led to a bi-selection framework for high-dimensional censored data. By integrating a composite penalty that targets both group structures and individual variables, and implementing a local coordinate descent algorithm, this approach achieves simultaneous identification of relevant variable groups and the key features within them. The resulting estimators exhibit oracle properties and strong performance in simulations, offering an alternative to proportional hazards frameworks when covariate effects accumulate additively on the hazard.
Variable selection techniques have also been extended to length-biased and interval-censored data, common in cohort and screening studies. A penalised expectation–maximisation algorithm with two-stage data augmentation overcomes the intractability of the likelihood function, yielding a stable and reliable method. The approach attains the oracle property and improves estimation accuracy over traditional conditional likelihood methods. Application to cancer screening trial data illustrates its capacity to pinpoint significant risk factors under complex sampling schemes.
Variable Selection and Estimation in Survival Analysis publication trend
The graph below shows the total number of articles in variable selection and estimation in survival analysis across all publications each year (not limited to Nature Index journals).
Technical terms
Censoring: A form of incomplete observation in which the exact event time is unknown but is known to lie above (right-censoring) or within (interval-censoring) certain bounds.
Left truncation: A bias that occurs when individuals enter a study only if their event time exceeds a threshold, requiring correction to avoid immortal time bias.
High-dimensional data: A setting where the number of covariates greatly exceeds the number of observations, necessitating specialised selection and regularisation methods.
Penalised regression: An estimation strategy that adds a penalty term (for example Lasso or group Lasso) to the likelihood to induce sparsity and prevent overfitting.
Length-biased data: Sampled observations whose inclusion probability is proportional to the event time, resulting in overrepresentation of longer durations.
References
- Penalized regression for left‐truncated and right‐censored survival data. Statistics in Medicine (2021).
- Bi-selection in the high-dimensional additive hazards regression model. Electronic Journal of Statistics (2021).
- Variable Selection for Length-Biased and Interval-Censored Failure Time Data. Mathematics (2023).
About these summaries
This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.
Turn complex research questions into confident strategic decisions
When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.
Benchmark your performance against global peers using robust, methodologically sound analysis.
Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.
Gain tailored, decision-ready recommendations aligned to your strategic priorities.
Talk to us to learn more about our data dashboards and bespoke strategy reports.
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.
Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:
Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.
Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.
Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.
Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.