Activation Function Strategies in Deep Learning Models

Summary

Activation functions are fundamental to the representational power of deep neural networks, introducing non-linearity that enables the modelling of complex patterns beyond linear relationships. Early approaches relied on simple threshold or sigmoid functions, which suffered from vanishing gradients and slow convergence in deep architectures. The introduction of the Rectified Linear Unit (ReLU) marked a turning point by offering computational simplicity, sparsity and robust gradient propagation, but it also led to issues such as “dying” neurons and unbounded activations. In response, researchers have developed a diverse array of smooth, non-monotonic and parametrically adaptable functions that combine the benefits of fast convergence, built-in regularisation and task-specific tuning. Current strategies include mathematically grounded functions with provable smoothness, hybrid constructions that regulate negative outputs, hardware-inspired forms reflecting neuromorphic principles and fully learnable families that evolve shape during gradient-based training. These advances have broadened the practical applicability of deep learning, improving accuracy and stability in computer vision, natural language processing, scientific modelling and reinforcement learning.

Research from Nature Portfolio

Recent studies have introduced a universal, trainable family of activation functions whose parameters evolve via standard optimisation to suit diverse tasks. In image classification with a VGG-style network on CIFAR-10, the function morphs into a variant resembling a known smooth form, matching or exceeding established benchmarks. In graph-based learning, it converges to identity to preserve spectral information, while in small-molecule quantification it adopts a hybrid that balances linear and saturating behaviour. For reinforcement learning, the method yields a novel shape that accelerates convergence. This adaptable paradigm demonstrates how a single parameterised form can replace multiple hand-crafted functions and deliver near-optimal performance across classification, regression and control problems.

Research from all publishers

A rigorous mathematical study of the Gaussian Error Linear Unit (GELU) has detailed its differentiability, boundedness and smoothness, and empirically confirmed superior performance to traditional and modern alternatives in residual convolutional networks on CIFAR-10, CIFAR-100 and STL-10. A novel non-monotonic function termed Smish leverages logarithmic scaling, sigmoid compression and a tanh operator to introduce negative-output regularisation; it outperforms leading activations on EfficientNet variants across CIFAR-10, MNIST and SVHN benchmarks. Additionally, a memristor-inspired activation function incorporates threshold and scaling parameters to create a smooth, non-monotonic response; experiments on multilayer perceptrons, ResNet, SqueezeNet and DenseNet architectures demonstrate faster convergence and improved accuracy compared with ReLU, Swish and related forms across multiple image-classification datasets.

Activation Function Strategies in Deep Learning Models publication trend

The graph below shows the total number of articles in activation function strategies in deep learning models across all publications each year (not limited to Nature Index journals).

Technical terms

Activation function: A mathematical operation applied to neuron inputs to introduce non-linearity, enabling networks to learn complex mappings between inputs and outputs.

Gaussian Error Linear Unit (GELU): A smooth activation defined as input multiplied by the Gaussian cumulative distribution, offering a balance between linear and non-linear regimes.

Smish: A hybrid activation combining logarithmic transformation, sigmoid compression and hyperbolic tangent scaling to introduce non-monotonicity and negative-output regulation.

ReLU-Memristor-like Activation Function (RMAF): A parametric function inspired by memristor circuits, featuring thresholding and scaling parameters that produce smooth, non-monotonic responses.

Universal Activation Function (UAF): A flexible family of functions with trainable parameters that evolve during model training into optimal shapes for diverse learning tasks.

References

  1. Universal activation function for machine learning. Scientific Reports (2021).
  2. Mathematical Analysis and Performance Evaluation of the GELU Activation Function in Deep Learning. Journal of Mathematics (2023).
  3. Smish: A Novel Activation Function for Deep Learning Methods. Electronics (2022).
  4. RMAF: Relu-Memristor-Like Activation Function for Deep Learning. IEEE Access (2020).

About these summaries

This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.

Nature Strategy Reports
Turn complex research questions into confident strategic decisions 

When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.

  • Benchmark your performance against global peers using robust, methodologically sound analysis.

  • Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.

  • Gain tailored, decision-ready recommendations aligned to your strategic priorities.

Talk to us to learn more about our data dashboards and bespoke strategy reports.

Nature Masterclasses
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.

Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:

  • Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.

  • Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.

  • Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.

Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.