Automatic Keyphrase Extraction Techniques in Natural Language Processing
Summary
Automatic keyphrase extraction aims to identify and rank representative words and multiword expressions that summarise a document’s content without manual annotation. Techniques fall broadly into unsupervised methods—relying on statistical distributions, linguistic patterns or graph-based representations—and supervised or deep learning approaches that leverage annotated corpora or pre-trained language models. Early work focused on term frequency–inverse document frequency and shallow co-occurrence graphs to discover salient terms. Graph-based algorithms such as TextRank and its variants model text as networks of words or phrases, while topic-based ranking methods incorporate latent semantic structures. Supervised sequence labelling uses feature-based classifiers or deep neural architectures to tag keyphrase boundaries directly, and the advent of transformer models has shifted attention towards contextual embeddings that capture fine-grained semantic relations. More recent hybrids blend contextual information at multiple granularities—sentences, paragraphs and entire documents—to improve domain independence and scalability. Evaluation typically employs precision, recall and F1-score on benchmark datasets spanning scientific articles, newswire and social media. Advances in centrality weighting, hierarchical topic discovery and knowledge-graph integration have driven performance gains. Practical applications range from scholarly search and metatagging to summarisation and recommendation systems, underscoring the global importance of automating high-level text comprehension across large and diverse corpora.
Research from Nature Portfolio
No recent Nature Portfolio content available.
Research from all publishers
Recent studies have advanced unsupervised methods by embedding richer contextual information into graph-based models, as exemplified by a hybrid technique that integrates sentence and paragraph embeddings into a hierarchical semantic graph to extract keyphrases and discover latent topics. This approach achieves significant improvements on both short and long documents. In parallel, deep learning research has enhanced sequence-labelling frameworks with pre-trained transformers: a centrality-weighted BERT model augments contextual embeddings with document-level relevance scores, yielding superior precision, recall and F1-score on scientific text corpora. Foundational work on transformer-based neural taggers has also demonstrated that adapting pre-trained language models to keyword identification tasks can closely match the performance of more resource-intensive supervised systems while requiring fewer annotations, broadening accessibility for low-resource domains.
Automatic Keyphrase Extraction Techniques in Natural Language Processing publication trend
The graph below shows the total number of articles in automatic keyphrase extraction techniques in natural language processing across all publications each year (not limited to Nature Index journals).
Technical terms
Unsupervised keyphrase extraction: Automatic identification of salient words or phrases without labelled data, using statistical, graph-based or linguistic cues.
Supervised sequence labelling: Learning to tag keyphrase boundaries using annotated corpora and feature-based or neural classifiers.
Graph-based model: Representation of text as a network of words or phrases, where centrality metrics rank candidate keyphrases.
Contextual embedding: Dense vector representation of words or sentences derived from deep language models that capture semantic and syntactic context.
Transformer: Neural architecture relying on self-attention mechanisms for contextualised language modelling.
Centrality weighting: Technique for integrating document-level relevance into token scoring, often via graph metrics or attention mechanisms.
References
- Contextual topic discovery using unsupervised keyphrase extraction and hierarchical semantic graph model. Journal of Big Data (2023).
- A Centrality-Weighted Bidirectional Encoder Representation from Transformers Model for Enhanced Sequence Labeling in Key Phrase Extraction from Scientific Texts. Big Data and Cognitive Computing (2024).
- TNT-KID: Transformer-based neural tagger for keyword identification. Natural Language Engineering (2021).
- Keyphrase Extraction Using Knowledge Graphs. Data Science and Engineering (2017).
About these summaries
This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.
Turn complex research questions into confident strategic decisions
When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.
Benchmark your performance against global peers using robust, methodologically sound analysis.
Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.
Gain tailored, decision-ready recommendations aligned to your strategic priorities.
Talk to us to learn more about our data dashboards and bespoke strategy reports.
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.
Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:
Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.
Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.
Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.
Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.