Statistical Dependence Measures in High-Dimensional Data
Summary
High-dimensional data have become ubiquitous across genomics, neuroscience, finance and other domains, characterised by a number of variables that can match or exceed the number of observations. Conventional correlation measures often fail to capture complex non-linear and multivariate interactions in such settings. Over the past decade, a variety of non-parametric dependence measures have emerged that extend beyond Pearson’s correlation and Spearman’s rho to detect subtle relationships among large variable sets. These include distance-based metrics, information-theoretic criteria, kernel-based tests and graph-theoretic approaches. Collectively, they aim to balance sensitivity to any form of association with computational feasibility in high dimensions. Applications range from inferring gene regulatory networks to identifying latent structure in imaging studies, while methodological challenges centre on reducing bias, improving power against diverse alternatives and scaling to modern data volumes.
Research from Nature Portfolio
An early seminal study presented an enhanced algorithm combining simulated annealing and genetic search to achieve exact maximal information coefficient calculations. This development overcame convergence issues of prior implementations, proved theoretical optimality and applied to million-pair gene expression profiles. By dramatically reducing estimation error and ensuring equitability across functional and non-functional relationships, it laid a rigorous foundation for information-based dependence discovery in large-scale biological datasets.
Research from all publishers
A user-friendly software platform implements distance correlation and its partial variant within a graphical interface, enabling simultaneous assessment of linear and non-linear associations in high-dimensional omics data. It supports one-to-one and one-to-all analyses, integrates Gaussian graphical modelling for conditional dependence estimation and streamlines network-driven biomarker discovery.
An improved algorithm for maximal information coefficient estimation introduces a back-search procedure on the grid partition axis, using a chi-square criterion to terminate optimisation. This approach yields more accurate partitions, preserves statistical power and reduces computation time. When applied to clustering of high-dimensional cancer genomics samples, it achieved superior separation of tumour and normal profiles compared with prior methods.
The multiscale graph correlation (MGC) framework combines k-nearest neighbour and kernel methods to detect dependencies across multiple scales. MGC demonstrates enhanced power in high-dimensional and non-linear scenarios, characterises the latent geometry underlying data relationships and remains computationally efficient. Applications in brain imaging and multi-omics integration showcase its ability to uncover interpretable dependency structure where traditional tests may lack sensitivity.
Statistical Dependence Measures in High-Dimensional Data publication trend
The graph below shows the total number of articles in statistical dependence measures in high-dimensional data across all publications each year (not limited to Nature Index journals).
Technical terms
High-dimensional data: Datasets in which the number of variables is comparable to or exceeds the number of observations, posing challenges for traditional statistical methods.
Distance correlation: A measure that quantifies both linear and non-linear dependence between random vectors by comparing pairwise distances in observation space.
Partial distance correlation: An extension of distance correlation that assesses the association between two variables while controlling for the influence of additional variables.
Maximal Information Coefficient (MIC): A statistic that identifies a wide range of associations by finding an optimal grid partition of the data to maximise mutual information.
Multiscale Graph Correlation (MGC): A test that integrates multiscale distance measures and graph-based neighbourhood analysis to detect and characterise dependencies in complex, high-dimensional datasets.
References
- Signed Distance Correlation (SiDCo): an online implementation of distance correlation and partial distance correlation for data-driven network analysis. Bioinformatics (2023).
- A Novel Algorithm for the Precise Calculation of the Maximal Information Coefficient. Scientific Reports (2014).
- A New Algorithm to Optimize Maximal Information Coefficient. PLOS ONE (2016).
About these summaries
This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.
Turn complex research questions into confident strategic decisions
When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.
Benchmark your performance against global peers using robust, methodologically sound analysis.
Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.
Gain tailored, decision-ready recommendations aligned to your strategic priorities.
Talk to us to learn more about our data dashboards and bespoke strategy reports.
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.
Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:
Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.
Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.
Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.
Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.