Summary

Adversarial machine learning studies the creation and defence against inputs—known as adversarial examples—that are intentionally perturbed to mislead trained models. Deep networks and other high-capacity learners exhibit fragile decision boundaries in high-dimensional spaces, enabling small, often imperceptible, perturbations to trigger misclassification or unsafe actions. Attacks are typically categorised by attacker knowledge—white-box attacks exploit full access to model parameters and gradients, whereas black-box attacks rely on input–output queries. Objectives range from targeted misdirection to untargeted disruption, with transferability of attacks threatening systems across platforms. The field has evolved into an arms race: novel attack algorithms prompt corresponding defences such as adversarial training, input transformations and certified robustness techniques. Addressing these vulnerabilities is critical for safety-critical applications—autonomous vehicles, healthcare diagnostics and security systems—and requires standardised threat models, benchmarks and interdisciplinary collaboration to ensure trust and reliability.

Research from Nature Portfolio

Investigations into medical-image diagnosis models have demonstrated the vulnerability of AI-CAD systems to adversarial images generated by generative adversarial networks. In one study, adversarial mammogram patches fooled a deep-learning breast-cancer detector in over two-thirds of cases that the model had originally classified correctly, while radiologists detected only a fraction of these manipulations, underscoring the need for integrated defence strategies and human-in-the-loop safeguards.

Adversarially trained networks, coupled with dual-batch-normalisation schemes, have been shown to enhance the clinical interpretability of saliency maps in X-ray, CT and MRI datasets without sacrificing diagnostic accuracy. When benchmarked on large external cohorts, these models matched the performance of standard networks while delivering more reliable and human-audit-friendly explanations of their decision pathways.

Research from all publishers

A comprehensive survey of adversarial attacks and defences in deep learning outlines the theoretical underpinnings of attack algorithms—gradient-based optimisations under norm constraints, surrogate-model black-box techniques and physical-world instantiations—alongside defensive measures including adversarial training, input randomisation and certified-radius guarantees. The review highlights open problems in adaptive attack-defence dynamics and calls for unified evaluation protocols.

In healthcare, reviews document the susceptibility of medical-image and wearable-sensor analytics to adversarial perturbations. Bespoke defence techniques such as model ensembling, uncertainty-aware sanitisation and privacy-preserving input reconstructions are surveyed, with an emphasis on developing regulatory-compliant robustness assessments and standardised benchmarks to guide deployment in safety-critical environments.

Adversarial Machine Learning publication trend

The graph below shows the total number of articles in adversarial machine learning across all publications each year (not limited to Nature Index journals).

Technical terms

Adversarial example: An input modified by subtle perturbations designed to mislead a machine-learning model while remaining indistinguishable to human perception.

White-box attack: An adversarial strategy in which the attacker has full access to the model’s architecture, parameters and gradients.

Black-box attack: An adversarial approach relying solely on input–output queries, without insight into the model’s internal structure.

Adversarial training: A defence technique that augments training data with adversarial examples to improve model robustness against perturbations.

Threat model: A formal specification of an adversary’s goals, knowledge assumptions and permissible actions within a given security scenario.

References

  1. Why deep-learning AIs are so easy to fool. Nature (2019).
  2. Adversarial Attacks and Defenses in Deep Learning. Engineering (2020).
  3. Secure and Robust Machine Learning for Healthcare: A Survey. IEEE Reviews in Biomedical Engineering (2021).
  4. A machine and human reader study on AI diagnosis model safety under attacks of adversarial images. Nature Communications (2021).
  5. Advancing diagnostic performance and clinical usability of neural networks via adversarial training and dual batch normalization. Nature Communications (2021).

About these summaries

This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.

Nature Strategy Reports
Turn complex research questions into confident strategic decisions 

When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.

  • Benchmark your performance against global peers using robust, methodologically sound analysis.

  • Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.

  • Gain tailored, decision-ready recommendations aligned to your strategic priorities.

Talk to us to learn more about our data dashboards and bespoke strategy reports.

Nature Masterclasses
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.

Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:

  • Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.

  • Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.

  • Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.

Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.