Summary

Positive and unlabeled (PU) learning addresses binary classification when only confirmed positive instances and a pool of unlabeled data are available. This paradigm combines elements of supervised and semi-supervised learning to reduce the cost and difficulty of obtaining negative labels. Typical approaches estimate the proportion of positives within the unlabeled set or identify a subset of reliable negatives, and then train a classifier on the labelled positives alongside these inferred negatives. Advances in the field have introduced generative frameworks, neighbourhood-based algorithms, tensor-network models and quantum-inspired strategies, each seeking improvements in accuracy, computational efficiency and robustness to distributional assumptions. Real-world applications span medical diagnosis, drug discovery, environmental risk mapping and web mining, where negative labels are scarce or prohibitively expensive. Research continues to refine loss functions, sampling schemes and evaluation metrics to address challenges such as class prior estimation bias, label noise and high dimensionality.

Research from Nature Portfolio

No recent Nature Portfolio content available.

Positive and Unlabeled Learning Techniques publication trend

The graph below shows the total number of articles in positive and unlabeled learning techniques across all publications each year (not limited to Nature Index journals).

Technical terms

Positive and Unlabeled (PU) Learning: A classification setting using only labelled positive examples and unlabelled data containing both positive and negative instances.

Class Prior: The true proportion of positive examples within the unlabeled set, essential for unbiased classifier training.

Reliable Negative Example: An unlabeled instance inferred with high confidence to belong to the negative class, used to complement positive labels in training.

Permutation Testing: A non-parametric method that assesses model performance by comparing results on true labels with those obtained under random label permutations.

Entropy Measure: A criterion for quantifying uncertainty or impurity in data partitions, commonly employed in decision-tree algorithms.

References

  1. Conditional generative positive and unlabeled learning. Expert Systems with Applications (2023).
  2. Positive unlabeled learning with tensor networks. Neurocomputing (2023).
  3. A Novel Classification Method: Neighborhood-Based Positive Unlabeled Learning Using Decision Tree (NPULUD). Entropy (2024).
  4. Leveraging permutation testing to assess confidence in positive-unlabeled learning applied to high-dimensional biological datasets. BMC Bioinformatics (2024).
  5. Positive Unlabeled Learning Selected Not At Random (PULSNAR): class proportion estimation without the selected completely at random assumption. PeerJ Computer Science (2024).
  6. A Quantum-Inspired Direct Learning Strategy for Positive and Unlabeled Data. International Journal of Computational Intelligence Systems (2023).

About these summaries

This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.

Nature Strategy Reports
Turn complex research questions into confident strategic decisions 

When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.

  • Benchmark your performance against global peers using robust, methodologically sound analysis.

  • Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.

  • Gain tailored, decision-ready recommendations aligned to your strategic priorities.

Talk to us to learn more about our data dashboards and bespoke strategy reports.

Nature Masterclasses
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.

Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:

  • Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.

  • Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.

  • Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.

Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.