Summary

Data mining and knowledge discovery refers to the systematic process of extracting useful patterns, trends and models from large and complex data sets. At its heart lies a cycle of problem definition, data preparation, model building and validation, followed by interpretation and deployment of results. Data may originate from diverse sources—relational databases, data warehouses, sensor streams or unstructured text—and typically require cleaning, transformation and dimensional‐reduction before analysis. Algorithms span supervised learning (classification, regression), unsupervised learning (clustering, association rule mining) and hybrid methods (ensemble learning, anomaly detection). Advances in computing power and algorithmic efficiency have made it possible to handle ever larger volumes of data, to discover non‐linear relationships via trees, neural networks or kernel methods, and to visualise multivariate structures through interactive plots and network representations. Practical applications range from healthcare diagnostics and fraud detection to supply‐chain optimisation, personalised marketing and scientific discovery, where the ultimate goal is to turn raw data into actionable insights that support decision‐making.

Research from Nature Portfolio

A recent study introduced a novel method for generating and combining basic probability assignments in uncertain information fusion. By integrating Mahalanobis and cosine similarities with belief entropy, the approach constructs evidence sources that resolve classical paradoxes in Dempster–Shafer theory. Implemented within an adversarial training loop, the method yields probability mass functions that converge rapidly and produce more reliable fusion outcomes across highly conflicting sensor inputs. Another contribution benchmarked the robustness of machine‐learning classifiers for viral genome analysis by simulating errors characteristic of high‐throughput sequencing platforms. The framework perturbs SARS-CoV-2 sequences under varying noise budgets and finds that certain embedding methods and classifiers maintain high accuracy despite realistic error rates. These insights guide selection of models for large‐scale genomic surveillance under imperfect data conditions.

Research from all publishers

A parallelisable multi‐subset self‐expressive model for subspace clustering partitions high‐dimensional data into smaller dictionaries, solving sparse representations in parallel. The resulting clusters match global structure while achieving orders‐of‐magnitude speed‐ups on benchmark image and text data sets. In environmental monitoring, ensemble learning of gradient boosting variants (XGBoost, LightGBM, CatBoost) combined with attention‐based CNN‐LSTM architectures has produced state‐of‐the‐art forecasts of harmful algal blooms. Bayesian hyperparameter optimisation and stacking improve prediction accuracy and stability across diverse coastal regions. In finance, a comprehensive review of transformer, GAN and graph‐neural‐network models for price forecasting highlights the superior capacity of attention mechanisms to capture long‐range dependencies, and recommends extending point forecasts to interval and density predictions to convey uncertainty more effectively in risk management applications.

Data Mining and Knowledge Discovery publication trend

The graph below shows the total number of articles in data mining and knowledge discovery across all publications each year (not limited to Nature Index journals).

Technical terms

Knowledge discovery (KDD): A multistage process of selecting, preprocessing, modelling and interpreting data to extract actionable patterns.

Supervised learning: Methods that infer a mapping from inputs to labelled outputs, including classification and regression.

Unsupervised learning: Techniques that identify inherent structures—clusters or association rules—without target labels.

Decision tree: A non‐parametric model that recursively partitions input space into homogeneous regions via binary splits.

Ensemble learning: Strategies that combine multiple models—bagging, boosting or stacking—to improve predictive performance.

Belief entropy: A measure of uncertainty in Dempster–Shafer evidence theory, generalising Shannon entropy for mass functions.

Embedding: A transformation of high‐dimensional data (e.g. sequences, graphs) into lower‐dimensional vectors that preserve salient features.

References

  1. A new basic probability assignment generation and combination method for conflict data fusion in the evidence theory. Scientific Reports (2023).
  2. Benchmarking machine learning robustness in Covid-19 genome sequence classification. Scientific Reports (2023).
  3. PMSSC: Parallelizable multi-subset based self-expressive model for subspace clustering. Computational Visual Media (2023).
  4. Ensemble Machine Learning of Gradient Boosting (XGBoost, LightGBM, CatBoost) and Attention-Based CNN-LSTM for Harmful Algal Blooms Forecasting. Toxins (2023).
  5. Deep learning models for price forecasting of financial time series: A review of recent advancements: 2020–2022. Wiley Interdisciplinary Reviews Data Mining and Knowledge Discovery (2023).

About these summaries

This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.

Nature Strategy Reports
Turn complex research questions into confident strategic decisions 

When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.

  • Benchmark your performance against global peers using robust, methodologically sound analysis.

  • Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.

  • Gain tailored, decision-ready recommendations aligned to your strategic priorities.

Talk to us to learn more about our data dashboards and bespoke strategy reports.

Nature Masterclasses
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.

Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:

  • Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.

  • Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.

  • Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.

Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.