Active Learning Strategies in Machine Learning Systems

Summary

Active learning encompasses a suite of iterative sampling methodologies that seek to maximise predictive performance while minimising the burden of manual annotation. By identifying and labelling only the most informative or uncertain instances, these strategies address the critical challenge of acquiring high‐quality labelled data in domains where annotation is costly or scarce. Central approaches include uncertainty sampling, where the model selects examples with maximal prediction ambiguity; query‐by‐committee, which exploits disagreements among multiple learners; and expected model change, which prioritises points that would induce the greatest update to the model’s parameters. Recent advances have fused traditional statistical criteria with deep architectures, enabling scalable pool‐based and batch‐mode workflows in vision and language tasks. Emerging trends further integrate auxiliary information—such as human cognitive signals or partial annotations—to refine query strategies and reduce labelling overhead. Across diverse applications from medical image segmentation to autonomous driving, active learning has demonstrated substantial gains in annotation efficiency, facilitating rapid deployment of robust systems under constrained budgets.

Research from Nature Portfolio

Innovative work has introduced neurally weighted sampling, wherein functional brain imaging guides the selection of training examples for object recognition. Leveraging human neural responses as soft constraints, this paradigm has yielded classifiers that align more closely with human perceptual features and achieve higher accuracy without additional neural data at inference. In parallel, a cascaded three‐dimensional network combined with an active learning loop has been applied to abdominal CT kidney segmentation. By iteratively correcting labels with a convolutional network and selectively querying the most uncertain regions, this approach halved manual labelling time while sustaining or improving segmentation accuracy across successive training stages.

Research from all publishers

A novel sample prioritisation framework draws on information‐value metrics inspired by reinforcement learning to navigate large unlabeled pools. This method balances exploitation of high‐density regions with exploration of rare but informative samples, demonstrating improved classification results in supervised scenarios. In semantic segmentation, partial annotation strategies have been proposed that exploit model uncertainty to guide human annotators to the most ambiguous or error‐prone image regions. This selective curation of ground truth labels markedly reduces annotation effort without compromising segmentation quality. Meanwhile, investigations into uncertainty measures for transformer models have revealed that conventional softmax‐based scores often target outliers rather than truly informative samples. By introducing a clipping heuristic to discard extreme probabilities, researchers have refined uncertainty sampling so as to focus on instances that genuinely reduce model uncertainty, restoring performance above random baselines.

Active Learning Strategies in Machine Learning Systems publication trend

The graph below shows the total number of articles in active learning strategies in machine learning systems across all publications each year (not limited to Nature Index journals).

Technical terms

Active learning: An iterative process in which a model queries an oracle to label the most informative unlabeled data points, aiming to maximise performance with minimal annotations.

Uncertainty sampling: A query strategy that selects instances for which the model’s prediction confidence is lowest, under the premise that labelling these will yield the greatest informational gain.

Query‐by‐committee: A technique employing an ensemble of models whose disagreement on unlabeled instances informs the choice of data to label next.

Batch mode: A sampling regime in which multiple instances are selected and labelled in each active learning iteration, as opposed to one‐by‐one selection.

Annotation efficiency: A measure of reduction in manual labelling effort achieved by employing selective sampling strategies rather than random sampling.

References

  1. Strategic data navigation: information value-based sample selection. Artificial Intelligence Review (2024).
  2. Partial annotations in active learning for semantic segmentation. Automation in Construction (2024).
  3. Comparing and Improving Active Learning Uncertainty Measures for Transformer Models by Discarding Outliers. Information Systems Frontiers (2024).
  4. Using human brain activity to guide machine learning. Scientific Reports (2018).
  5. Active learning for accuracy enhancement of semantic segmentation with CNN-corrected label curations: Evaluation on kidney segmentation in abdominal CT. Scientific Reports (2020).
  6. A Survey on Deep Active Learning: Recent Advances and New Frontiers. IEEE Transactions on Neural Networks and Learning Systems (2025).

About these summaries

This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.

Nature Strategy Reports
Turn complex research questions into confident strategic decisions 

When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.

  • Benchmark your performance against global peers using robust, methodologically sound analysis.

  • Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.

  • Gain tailored, decision-ready recommendations aligned to your strategic priorities.

Talk to us to learn more about our data dashboards and bespoke strategy reports.

Nature Masterclasses
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.

Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:

  • Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.

  • Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.

  • Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.

Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.