Author Name Disambiguation in Scholarly Communications
Summary
Author name disambiguation addresses the critical challenge of correctly attributing scholarly works to individual researchers despite the prevalence of identical or similar names. Without reliable disambiguation, bibliometric analyses, institutional evaluations and literature searches can be significantly distorted by splitting errors—where one individual is treated as several—and lumping errors—where distinct individuals are conflated. Advances in metadata enrichment, algorithmic sophistication and authority control infrastructures have driven progress in this field. Contemporary approaches draw upon supervised and unsupervised machine learning, network and graph theory, natural language processing and persistent identifiers to improve precision and recall. The global research community has adopted layered solutions, combining rule-based heuristics with data-driven models, while emphasising transparency, scalability and ongoing validation. Practical deployments in major digital libraries and aggregators now support researcher profiling, grant reporting and discovery services. Yet challenges remain in handling sparse metadata contexts, rapidly growing publication volumes and culturally diverse naming conventions. Emerging trends include the integration of knowledge graphs, multilingual name processing and ethnicity-sensitive models, all underscoring the continued interdisciplinary importance of author name disambiguation for reliable scholarly communication and evaluation on a worldwide scale.
Research from Nature Portfolio
No recent Nature Portfolio content available.
Author Name Disambiguation in Scholarly Communications publication trend
The graph below shows the total number of articles in author name disambiguation in scholarly communications across all publications each year (not limited to Nature Index journals).
Technical terms
Homonymy: Occurrence of identical names belonging to different individuals, leading to potential conflation.
Lumping error: Incorrect merging of publications by distinct authors under a single name entity.
Splitting error: Erroneous separation of publications by one author into multiple entities due to name variants.
Knowledge Graph Embedding: Representation of nodes and relationships in a continuous vector space to capture structural and attribute information.
Agglomerative Clustering: Hierarchical clustering method that iteratively merges the most similar clusters based on defined linkage criteria.
ORCID: A non-proprietary persistent digital identifier that uniquely distinguishes researchers across publications and platforms.
References
- PubMed Computed Authors in 2024: an open resource of disambiguated author names in biomedical literature. Bioinformatics (2024).
- Graph-based methods for Author Name Disambiguation: a survey. PeerJ Computer Science (2023).
- A knowledge graph embeddings based approach for author name disambiguation using literals. Scientometrics (2022).
- Ethnicity‐based name partitioning for author name disambiguation using supervised machine learning. Journal of the Association for Information Science and Technology (2021).
About these summaries
This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.
Turn complex research questions into confident strategic decisions
When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.
Benchmark your performance against global peers using robust, methodologically sound analysis.
Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.
Gain tailored, decision-ready recommendations aligned to your strategic priorities.
Talk to us to learn more about our data dashboards and bespoke strategy reports.
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.
Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:
Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.
Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.
Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.
Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.