Deep Learning Techniques for Oracle Bone Inscriptions Recognition
Summary
Oracle bone inscriptions represent the earliest known corpus of Chinese writing and remain a critical resource for understanding ancient language, culture and history. Recent advances in deep learning have transformed the automatic recognition of these challenging glyphs, which exhibit considerable noise, fragmentation and stylistic variation after millennia underground. Convolutional neural networks have been adapted with attention modules and contextual transformer blocks to enhance robustness against blur, occlusion and damage. Meanwhile, one‐shot and few‐shot learning frameworks, such as Siamese networks, address the chronic scarcity and class imbalance intrinsic to oracle datasets by learning similarity metrics directly from image pairs. Transformer‐based architectures, including Swin‐Transformer variants, further capture long‐range dependencies and fine‐grained morphological features. Successive contributions have also focused on creating benchmark datasets—ranging from MNIST‐compatible collections of binary oracle characters to large-scale multilingual corpora—and on applying sophisticated data-augmentation pipelines to simulate aged inscriptions. Collectively, these approaches have steadily improved top-1 and top-5 recognition accuracies, enabled open-set classification of unknown glyphs and laid the groundwork for integrated pipelines that detect, segment and interpret oracle characters in both 2D rubbings and 3D bone models.
Research from Nature Portfolio
Recent studies have introduced a Siamese similarity network tailored to ancient character recognition, directly learning pairwise similarity to support one-shot classification. This network employs a multi-scale fusion backbone and an embedded soft similarity contrast loss to enhance discrimination between visually similar glyphs while mitigating over-fitting. A cumulative class prototype strategy refines class representations and enables the model to reject unknown categories gracefully. Experimental evidence demonstrates that this approach achieves state-of-the-art performance against traditional convolutional architectures and classic one-shot learning methods, offering a scalable solution for recognising newly discovered oracle characters with minimal training examples.
Research from all publishers
Several publications have advanced dataset creation and model design for oracle bone recognition. A new Oracle-MNIST benchmark provides over 30,000 grayscale images of 30 glyph classes formatted for direct compatibility with existing classifiers, challenging models to cope with extreme noise and diverse writing styles. An improved Inception-v3 network integrates contextual transformer blocks and a convolutional attention module to boost top-1, top-3 and top-5 accuracies on two standard oracle datasets, demonstrating superior resilience to blurred and mutilated inscriptions compared with AlexNet and VGG-19 baselines. In parallel, a large-scale ancient character dataset of over 9,000 distinct classes enables a Swin-Transformer-based model with linear embedding and CoT blocks to attain top-one accuracy above 87% and top-five accuracy near 96%, aided by adaptive data-enhancement and resampling strategies that capture morphological nuances across thousands of archaeological images.
Deep Learning Techniques for Oracle Bone Inscriptions Recognition publication trend
The graph below shows the total number of articles in deep learning techniques for oracle bone inscriptions recognition across all publications each year (not limited to Nature Index journals).
Technical terms
Oracle bone inscriptions: Ancient Chinese characters inscribed on turtle shells or animal bones dating to the Shang dynasty, often degraded by age.
Convolutional neural network (CNN): A deep learning model employing convolutional layers to extract spatial hierarchies of features from images.
Siamese similarity network: A neural architecture that processes pairs of inputs through twin branches to learn a distance metric for one-shot or few-shot classification.
One-shot learning: A learning paradigm in which a model recognises new classes from a single example by leveraging similarity measures and prior knowledge.
Transformer block: A self-attention module that captures long-range dependencies in data, adapted here to focus on fine-grained character morphology.
Data augmentation: Techniques that expand and diversify training data by applying transformations such as rotation, scaling, noise injection or symmetry operations.
References
- A dataset of oracle characters for benchmarking machine learning algorithms. Scientific Data (2024).
- Ancient Chinese Character Recognition with Improved Swin-Transformer and Flexible Data Enhancement Strategies. Sensors (2024).
- An Improved Neural Network Model Based on Inception‐v3 for Oracle Bone Inscription Character Recognition. Scientific Programming (2022).
- One shot ancient character recognition with siamese similarity network. Scientific Reports (2022).
About these summaries
This Nature Research Intelligence Topic summary is created with the cited references and a large language model. We take care to ground generated text with facts, and have systems in place to gain human feedback on the overall quality of the process in line with our AI principles. We strive to create accurate and useful summaries for people unfamiliar with the research topic and that supports this goal. These pages are a beta release and will be updated as we learn how best to help people gain value from a research topic summary.
Turn complex research questions into confident strategic decisions
When you're under pressure to set direction, justify investment, or understand your competitive position, you need more than raw data — you need trusted insights you can act on.
Benchmark your performance against global peers using robust, methodologically sound analysis.
Combine quantitative metrics with qualitative expert insight to uncover strengths, gaps and emerging opportunities.
Gain tailored, decision-ready recommendations aligned to your strategic priorities.
Talk to us to learn more about our data dashboards and bespoke strategy reports.
Grow research skills, confidence and careers with training built for every stage of the research lifecycle.
Developed with Nature Portfolio journal Editors and internationally renowned experts. Discover three ways to learn:
Self-paced, online courses in convenient bite-sized units, covering key skills across scientific writing, publishing, grant writing, data analysis, and more.
Expert trainer-led workshops with hands-on exercises and real-time feedback across core research skills, delivered via interactive group sessions.
Editor-led workshops combining core principles in writing and publishing, personalised 1:1 feedback from Nature Portfolio Editors and hands-on exercises.
Explore course catalogues and workshop agendas, enquire about the options or request institutional pricing.