paper-with-me

홈 › Papers

Mitigating Long-tail Distribution in Oracle Bone Inscriptions: Dataset, Model, and Benchmark

2025-04-13 · Jinhao Li, Zijian Chen, Runze Jiang, Tingzhu Chen, Changbo Wang, Guangtao Zhai

The oracle bone inscription (OBI) recognition plays a significant role in understanding the history and culture of ancient China. However, the existing OBI datasets suffer from a long-tail distribution problem, leading to biased performance of OBI recognition models across majority and minority classes. With recent advancements in generative models, OBI synthesis-based data augmentation has become a promising avenue to expand the sample size of minority classes. Unfortunately, current OBI datasets lack large-scale structure-aligned image pairs for generative model training. To address these problems, we first present the Oracle-P15K, a structure-aligned OBI dataset for OBI generation and denoising, consisting of 14,542 images infused with domain knowledge from OBI experts. Second, we propose a diffusion model-based pseudo OBI generator, called OBIDiff, to achieve realistic and controllable OBI generation. Given a clean glyph image and a target rubbing-style image, it can effectively transfer the noise style of the original rubbing to the glyph image. Extensive experiments on OBI downstream tasks and user preference studies show the effectiveness of the proposed Oracle-P15K dataset and demonstrate that OBIDiff can accurately preserve inherent glyph structures while transferring authentic rubbing styles effectively.

📄 PDF Abstract BibTeX arXiv:2504.09555

Code (0)

등록된 구현이 없습니다.

Tasks

Data AugmentationDenoising

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

OracleAnalyser: Analysing Implicit Semantics of Oracle Bone Scripts through MLLMs with Post-training

2026-06-24 · Zijia Song, Yelin Wang, Zhengyi Ma, Zitong Yu 외 arxiv

With the advancement of artificial intelligence, research on oracle bone scripts has entered a new era. However, existing methods and benchmarks remain largely confined to recognition tasks, overlooking the equally cruci…

Diff-Oracle: Deciphering Oracle Bone Scripts with Controllable Diffusion Model

2023-12-21 · Jing Li, Qiu-Feng Wang, Siyuan Wang, Rui Zhang 외

Deciphering oracle bone scripts plays an important role in Chinese archaeology and philology. However, a significant challenge remains due to the scarcity of oracle character images. To overcome this issue, we propose Di…

Image GenerationImage-to-Image Translation

An open dataset for oracle bone script recognition and decipherment

2024-01-27 · Pengjie Wang, Kaile Zhang, Xinyu Wang, Shengwei Han 외

Oracle bone script, one of the earliest known forms of ancient Chinese writing, presents invaluable research materials for scholars studying the humanities and geography of the Shang Dynasty, dating back 3,000 years. The…

Decipherment

Long-Tailed Partial Label Learning via Dynamic Rebalancing

2023-02-10 · Feng Hong, Jiangchao Yao, Zhihan Zhou, Ya zhang 외

Real-world data usually couples the label ambiguity and heavy imbalance, challenging the algorithmic robustness of partial label learning (PLL) and long-tailed learning (LT). The straightforward combination of LT and PLL…

Partial Label Learning

Boosting Long-tailed Object Detection via Step-wise Learning on Smooth-tail Data

2023-05-22 · ICCV 2023 1 · Na Dong, Yongqiang Zhang, Mingli Ding, Gim Hee Lee

Real-world data tends to follow a long-tailed distribution, where the class imbalance results in dominance of the head classes during training. In this paper, we propose a frustratingly simple but effective step-wise lea…

AllLong-tailed Object Detectionobject-detectionObject Detection