Data Imputation Techniques for Hydrological Time Series
Summary
Hydrological time series often exhibit missing values due to instrument failure, maintenance gaps or transmission errors. To maintain dataset continuity and support robust analysis of streamflow, precipitation and water‐quality trends, a variety of imputation techniques have been developed. Simple approaches such as linear interpolation and nearest‐neighbour substitution perform well for short gaps but may bias longer sequences. Geostatistical methods, including inverse distance weighting and spatio‐temporal kriging, exploit spatial coherence among monitoring sites to reconstruct missing records. Machine learning algorithms such as random forests, neural networks and multiple imputation by chained equations adaptively learn nonlinear relationships among variables and can handle high percentages of missing data. Debiasing of climate reanalysis products offers an additional source of auxiliary information for extended gaps, while hydrological models calibrated on observed catchment responses can generate physically consistent estimates of unobserved flows. Hybrid frameworks that sequence or combine these methods—selecting the optimal approach for each gap based on its length, seasonality and data context—have demonstrated improved accuracy and reduced uncertainty. Performance is assessed by metrics such as the Nash–Sutcliffe efficiency, root‐mean‐square error and Kling–Gupta efficiency, and bias‐correction steps are often applied to preserve variance and long‐term trends. Reliable imputation underpins flood forecasting, water‐resource planning and climate‐change impact assessments across diverse regions worldwide.
Research from Nature Portfolio
No recent Nature Portfolio content available.
Data Imputation Techniques for Hydrological Time Series publication trend
The graph below shows the total number of articles in data imputation techniques for hydrological time series across all publications each year (not limited to Nature Index journals).
Technical terms
Linear interpolation: A method that estimates missing values by drawing straight lines between consecutive observed points.
Spatio‐temporal kriging: A geostatistical technique that predicts missing observations by modelling spatial and temporal correlations simultaneously.
Reanalysis data: Retrospective gridded climate products combining observations and numerical models to provide continuous estimates of meteorological variables.
Random forest: A machine learning ensemble of decision trees used for imputing values based on patterns in multiple predictor variables.
Multiple imputation by chained equations (MICE): An iterative technique that fills missing data by generating several plausible values and pooling results across imputations.
Nash–Sutcliffe efficiency (NSE): A normalized statistic that assesses the predictive power of imputed or modelled series relative to observed data.
Kling–Gupta efficiency (KGE): A performance metric that combines correlation, bias and variability to evaluate hydrological reconstructions.
References
- Gap-free 16-year (2005–2020) sub-diurnal surface meteorological observations across Florida. Scientific Data (2023).
- Filling gaps in urban temperature observations by debiasing ERA5 reanalysis data. Urban Climate (2024).
- Interpolation in Time Series: An Introductive Overview of Existing Methods, Their Performance Criteria and Uncertainty Assessment. Water (2017).
- Estimating extremely large amounts of missing precipitation data. Journal of Hydroinformatics (2020).
- How good are hydrological models for gap-filling streamflow data?. Hydrology and Earth System Sciences (2018).
Turn complex research questions into confident strategic decisions
When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.
Benchmark your performance against global peers using robust, methodologically sound analysis.
Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.
Gain tailored, decision-ready recommendations aligned to your strategic priorities.
Talk to us to learn more about our data dashboards and bespoke strategy reports.
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.
Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:
Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.
Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.
Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.
Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.