Automatic Speech Disfluency Detection Using Deep Learning Techniques

Summary

Automatic speech disfluency detection has become a vital component of modern speech processing, with applications ranging from clinical assessment of stuttering to the enhancement of conversational agents and transcription services. Deep learning has transformed this domain by enabling end-to-end frameworks that learn discriminative representations directly from raw audio or intermediate speech embeddings. Early approaches relied on hand-crafted acoustic features and classical classifiers, but recent advances exploit neural architectures such as bidirectional recurrent networks, convolutional layers and attention-driven Transformers. These models can distinguish between fluent speech and a variety of disfluent behaviours—such as sound or syllable repetitions, prolongations and filled pauses—while accommodating variable segment lengths, multilingual data and limited annotated corpora. By integrating self-supervised pre-training, multi-feature fusion and context-aware sequence modelling, contemporary systems achieve robust detection performance, contributing to improved therapeutic feedback, richer metrics for speech-technology benchmarks and more natural human–machine interaction worldwide.

Research from Nature Portfolio

No recent Nature Portfolio content available.

Research from all publishers

Among non-Portfolio publications, several deep learning strategies have demonstrated marked improvements in disfluency detection. One study introduced an attention-enhanced model that fuses multiple acoustic descriptors—including pitch, temporal cues and spectral features—with bidirectional long short-term memory units. Spatial and temporal attention mechanisms prioritise salient segments, yielding substantial gains in F1 scores across standard stuttering databases. Another work proposed “TranStutter”, a convolution-free Transformer that employs multi-head self-attention and positional encoding to classify stuttering events. Evaluated on podcast-derived and clinical interview datasets, this approach achieved accuracies exceeding 85 per cent, illustrating the capacity of pure attention models to capture nuanced temporal patterns without convolutional preprocessing. A third line of research leverages self-supervised speech representations from wav2vec2.0, integrating pre-trained embeddings with combined convolutional and Transformer layers to detect disfluencies in multiple languages and across variable-length utterances. This method not only reduces reliance on extensive manual annotation but also demonstrates scalability to low-resource languages, achieving consistent detection rates in English and Chinese corpora and showing promise for deployment in diverse linguistic environments.

Automatic Speech Disfluency Detection Using Deep Learning Techniques publication trend

The graph below shows the total number of articles in automatic speech disfluency detection using deep learning techniques across all publications each year (not limited to Nature Index journals).

Technical terms

Speech disfluency: Interruptions or irregularities in the normal flow of speech, including hesitations, repetitions and prolongations.

Attention mechanism: A neural network component that weighs the importance of different input segments to focus processing on relevant features.

Transformer: A deep learning architecture based on self-attention layers and positional encoding, enabling parallel sequence processing without recurrence.

wav2vec2.0: A self-supervised model that learns contextualised speech representations from large unlabelled audio datasets for downstream tasks.

Bidirectional Long Short-Term Memory (BiLSTM): A recurrent network that processes sequences in both forward and backward directions to capture past and future context.

References

  1. A novel attention model across heterogeneous features for stuttering event detection. Expert Systems with Applications (2024).
  2. TranStutter: A Convolution-Free Transformer-Based Deep Learning Method to Classify Stuttered Speech Using 2D Mel-Spectrogram Visualization and Attention-Based Feature Representation. Sensors (2023).
  3. Automatic Speech Disfluency Detection Using wav2vec2.0 for Different Languages with Variable Lengths. Applied Sciences (2023).

About these summaries

This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.

Nature Strategy Reports
Turn complex research questions into confident strategic decisions 

When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.

  • Benchmark your performance against global peers using robust, methodologically sound analysis.

  • Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.

  • Gain tailored, decision-ready recommendations aligned to your strategic priorities.

Talk to us to learn more about our data dashboards and bespoke strategy reports.

Nature Masterclasses
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.

Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:

  • Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.

  • Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.

  • Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.

Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.