Causal Inference Methods in Observational Studies
Summary
Causal inference in observational settings seeks to estimate the effect of exposures, treatments or interventions on outcomes in the absence of random assignment. Unlike experimental designs, observational studies must address confounding arising from systematic differences between exposed and unexposed groups. A range of approaches has been developed to mimic randomisation, including propensity-score methods, inverse probability weighting, matching and stratification. At the same time, graphical frameworks such as directed acyclic graphs (DAGs) and structural causal models (SCMs) provide a formal language for articulating assumptions, identifying minimal adjustment sets and diagnosing potential biases. Counterfactual reasoning underpins these methods by defining quantities such as the average treatment effect that compare hypothetical outcomes under different exposure regimes. Recent advances have focused on sensitivity analysis to assess the impact of unmeasured confounders, on data-driven variable selection guided by causal criteria, and on machine-learning-based propensity-score estimation while maintaining interpretability. Applications span epidemiology, economics, social science and public policy, where the accurate estimation of causal effects informs healthcare decisions, resource allocation and regulatory interventions. By combining statistical rigour with clear causal assumptions, modern methods in observational causal inference offer robust tools for generating evidence when trials are impractical or unethical.
Research from Nature Portfolio
No recent Nature Portfolio content available.
Research from all publishers
Recent studies have sought to disambiguate the assumptions required for each stage of causal analysis, from discovery to formal inference. A comprehensive review has mapped a variety of causal concepts onto different levels of the causal hierarchy, examining which graphical and parametric assumptions are needed to move from association to policy-relevant causal effects, counterfactual queries and mediation analysis. This work emphasises the importance of explicitly stating and testing the assumptions that underpin each causal claim.
The use of directed acyclic graphs in applied health research has been surveyed to identify inconsistencies in reporting and practice. Analyses revealed substantial variation in how target estimands, adjustment sets and unobserved confounders are depicted, leading to recommendations for improving transparency, DAG standardisation and documentation of underlying causal assumptions in published studies.
New principles for confounder selection have been formulated to guide covariate control when full causal diagrams are unavailable. A practical strategy is proposed: adjust for any variable that is a cause of the exposure or the outcome, exclude instrumental variables, and include proxies for unmeasured common causes. This approach aligns theoretical developments with statistical covariate-selection methods and aims to simplify real-world implementation.
Causal Inference Methods in Observational Studies publication trend
The graph below shows the total number of articles in causal inference methods in observational studies across all publications each year (not limited to Nature Index journals).
Technical terms
Observational study: A study in which the investigator does not control assignment to exposure or treatment, relying on naturally occurring variation.
Confounding: A biasing effect that arises when a third variable influences both the exposure and the outcome.
Propensity score: The probability of receiving a particular exposure conditional on observed covariates, used to balance treated and untreated groups.
Inverse probability of treatment weighting (IPTW): A method that weights each subject by the inverse of their propensity score to create a synthetic population in which exposure is independent of measured covariates.
Directed acyclic graph (DAG): A graphical representation of causal relations among variables, consisting of nodes and directed edges without feedback loops.
Structural causal model (SCM): A formal system of equations and graphical structures that defines causal mechanisms and counterfactual quantities.
References
- Causal inference in statistics: An overview. Statistics Surveys (2009).
- Sensitivity Analysis Without Assumptions. Epidemiology (2016).
- Disentangling causality: assumptions in causal discovery and inference. Artificial Intelligence Review (2023).
- Use of directed acyclic graphs (DAGs) to identify confounders in applied health research: review and recommendations. International Journal of Epidemiology (2020).
- Principles of confounder selection. European Journal of Epidemiology (2019).
About these summaries
This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.
Turn complex research questions into confident strategic decisions
When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.
Benchmark your performance against global peers using robust, methodologically sound analysis.
Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.
Gain tailored, decision-ready recommendations aligned to your strategic priorities.
Talk to us to learn more about our data dashboards and bespoke strategy reports.
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.
Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:
Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.
Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.
Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.
Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.