Phylogenetic Data Analysis and Inference Techniques
Summary
Phylogenetic data analysis encompasses an array of computational and statistical methods designed to reconstruct the evolutionary relationships among organisms, genes or genomes. Central to this endeavour is the comparison of homologous characters—typically nucleotide or amino acid sequences—across taxa to infer branching patterns and divergence times. Modern inference techniques employ both concatenation and coalescent frameworks, enabling researchers to combine data from multiple loci while accounting for gene-tree/species-tree discordance. Advances in model development have addressed biases arising from compositional heterogeneity, rate variation among sites and long-branch attraction, with site-heterogeneous models such as CAT and CAT-GTR mitigating systematic artefacts. Equally important are methods for optimising sequence alignments and trimming uninformative or misleading regions, as well as emerging approaches that exploit synteny and rare genomic changes to resolve deep phylogenetic nodes. Together, these innovations have enhanced our capacity to generate robust evolutionary hypotheses, illuminating patterns of biodiversity, informing conservation priorities and underpinning studies of trait evolution across the tree of life.
Research from Nature Portfolio
Recent large-scale phylogenomic reconstruction of prokaryotic diversity has yielded a reference tree encompassing over ten thousand bacterial and archaeal genomes. By integrating hundreds of conserved marker genes and evaluating the impact of taxon sampling, substitution heterogeneity and saturation, this work has revealed a closer evolutionary proximity between Archaea and Bacteria than previously inferred from ribosomal proteins alone. Robustness tests—including varied site-sampling schemes and exclusion of candidate phyla radiation taxa—demonstrate the stability of domain-level relationships under complex models, setting a new standard for microbial evolutionary frameworks.
Phylogenetic Data Analysis and Inference Techniques publication trend
The graph below shows the total number of articles in phylogenetic data analysis and inference techniques across all publications each year (not limited to Nature Index journals).
Technical terms
Phylogenetic inference: The process of reconstructing evolutionary relationships among taxa using genetic or morphological data and statistical models.
Multiple sequence alignment: A procedure that arranges sequences of DNA, RNA or protein to identify regions of similarity and infer homology.
Site-heterogeneous model: A class of evolutionary models that allows different sites in an alignment to follow distinct substitution processes or equilibrium frequencies.
Long-branch attraction: A phylogenetic artefact in which rapidly evolving lineages are erroneously inferred as closely related due to convergent substitutions.
Compositional heterogeneity: Variation in nucleotide or amino-acid frequencies across sites or lineages, which can bias standard substitution models.
Synteny: The preservation of gene order on chromosomes among different species, used as an additional character for phylogenetic reconstruction.
Parsimony-informative site: A position in an alignment at which at least two character states each occur in at least two sequences, contributing to tree resolution.
References
- A Bayesian Mixture Model for Across-Site Heterogeneities in the Amino-Acid Replacement Process. Molecular Biology and Evolution (2004).
- Phylogenomics of 10,575 genomes reveals evolutionary proximity between domains Bacteria and Archaea. Nature Communications (2019).
- The promise and pitfalls of synteny in phylogenomics. PLOS Biology (2024).
- Compositionally Constrained Sites Drive Long-Branch Attraction. Systematic Biology (2023).
- ClipKIT: A multiple sequence alignment trimming software for accurate phylogenomic inference. PLOS Biology (2020).
About these summaries
This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.
Turn complex research questions into confident strategic decisions
When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.
Benchmark your performance against global peers using robust, methodologically sound analysis.
Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.
Gain tailored, decision-ready recommendations aligned to your strategic priorities.
Talk to us to learn more about our data dashboards and bespoke strategy reports.
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.
Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:
Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.
Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.
Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.
Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.