Summary

Coding, information theory and compression provide a unified framework for quantifying, transmitting and summarising data. At its heart lies the concept of entropy, which measures the average uncertainty or information content of a probability distribution. Source coding exploits this by assigning short codewords to frequent symbols (Huffman or arithmetic coding), achieving near‐optimal average length. Channel coding uses structured redundancy (LDPC, turbo or polar codes) to approach Shannon capacity over noisy links. Universal coding methods, including two‐part codes and normalized maximum likelihood, extend these ideas when the true source distribution is unknown. Beyond statistical approaches, algorithmic information theory (Kolmogorov complexity) quantifies the information in individual strings via the length of the shortest generating programme. The Minimum Description Length principle unites these views by selecting models that minimise the total code‐length of model plus data, yielding a powerful paradigm for inference, clustering and model selection. In practical settings, lossless compression (ZIP, PNG) preserves all information under statistical or dictionary codes, while lossy schemes (JPEG, MP3) exploit perceptual irrelevance to achieve higher ratios. Recent advances extend classical rate–distortion theory into machine learning, simulation sciences and high‐dimensional data, guiding the design of efficient encoders and reliable communication protocols across diverse applications.

Research from Nature Portfolio

New analyses of discrete input–output mappings have shown that, under minimal assumptions, the probability of a randomly sampled input producing a given output decays exponentially with the output’s approximated Kolmogorov complexity. This “simplicity bias” was rigorously bounded and empirically validated in domains ranging from RNA folding to financial models, revealing a unifying principle across biological and engineered systems. Separately, an information‐theoretic approach to atmospheric and climate data defined “real information content” bits by identifying and rounding away noise‐level precision. Coupling this quantisation with standard lossless compressors achieved over 17× reduction on 64-bit floats while preserving 99 % of meaningful content—and more than 60× when spatio‐temporal correlations were exploited—offering a data‐compression Turing test to balance fidelity and compressibility.

Coding, Information Theory and Compression publication trend

The graph below shows the total number of articles in coding, information theory and compression across all publications each year (not limited to Nature Index journals).

Technical terms

Entropy: The expected information content of a random variable, defined as the average of –log p(x) over its distribution.

Prefix code: A set of codewords in which no codeword is a prefix of another, enabling instantaneous decoding.

Kolmogorov complexity: The length of the shortest universal-Turing‐machine programme that generates a given string; a measure of algorithmic information.

Normalized maximum likelihood (NML): A universal coding distribution formed by normalising maximum likelihoods over all sequences, minimising worst‐case regret.

Minimum Description Length (MDL) principle: A model selection criterion that chooses the hypothesis minimising the sum of model code-length and data code-length given the model.

References

  1. Input–output maps are strongly biased towards simple outputs. Nature Communications (2018).
  2. Compressing atmospheric data into its real information content. Nature Computational Science (2021).
  3. Toward a Simulation Model Complexity Measure. Information (2023).

About these summaries

This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.

Nature Strategy Reports
Turn complex research questions into confident strategic decisions 

When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.

  • Benchmark your performance against global peers using robust, methodologically sound analysis.

  • Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.

  • Gain tailored, decision-ready recommendations aligned to your strategic priorities.

Talk to us to learn more about our data dashboards and bespoke strategy reports.

Nature Masterclasses
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.

Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:

  • Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.

  • Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.

  • Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.

Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.