paper-with-me

홈 › Papers

InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling

2025-08-12 · Xiaolei Diao, Zhihan Zhou, Lida Shi, Ting Wang, Ruihua Qi, Hao Xu, Daqian Shi arxiv

Constructing historical language models (LMs) plays a crucial role in aiding archaeological provenance studies and understanding ancient cultures. However, existing resources present major challenges for training effective LMs on historical texts. First, the scarcity of historical language samples renders unsupervised learning approaches based on large text corpora highly inefficient, hindering effective pre-training. Moreover, due to the considerable temporal gap and complex evolution of ancient scripts, the absence of comprehensive character encoding schemes limits the digitization and computational processing of ancient texts, particularly in early Chinese writing. To address these challenges, we introduce InteChar, a unified and extensible character list that integrates unencoded oracle bone characters with traditional and modern Chinese. InteChar enables consistent digitization and representation of historical texts, providing a foundation for robust modeling of ancient scripts. To evaluate the effectiveness of InteChar, we construct the Oracle Corpus Set (OracleCS), an ancient Chinese corpus that combines expert-annotated samples with LLM-assisted data augmentation, centered on Chinese oracle bone inscriptions. Extensive experiments show that models trained with InteChar on OracleCS achieve substantial improvements across various historical language understanding tasks, confirming the effectiveness of our approach and establishing a solid foundation for future research in ancient Chinese NLP.

📄 PDF Abstract BibTeX arXiv:2508.15791

Code (0)

등록된 구현이 없습니다.

Tasks

Data Augmentation

Similar Papers 제목 키워드 기반

OracleSage: Towards Unified Visual-Linguistic Understanding of Oracle Bone Scripts through Cross-Modal Knowledge Fusion

2024-11-26 · Hanqi Jiang, Yi Pan, JunHao Chen, Zhengliang Liu 외

Oracle bone script (OBS), as China's earliest mature writing system, present significant challenges in automatic recognition due to their complex pictographic structures and divergence from modern Chinese characters. We …

An open dataset for the evolution of oracle bone characters: EVOBC

2024-01-23 · Haisu Guan, Jinpeng Wan, Yuliang Liu, Pengjie Wang 외

The earliest extant Chinese characters originate from oracle bone inscriptions, which are closely related to other East Asian languages. These inscriptions hold immense value for anthropology and archaeology. However, de…

Decipherment

Diff-Oracle: Deciphering Oracle Bone Scripts with Controllable Diffusion Model

2023-12-21 · Jing Li, Qiu-Feng Wang, Siyuan Wang, Rui Zhang 외

Deciphering oracle bone scripts plays an important role in Chinese archaeology and philology. However, a significant challenge remains due to the scarcity of oracle character images. To overcome this issue, we propose Di…

Image GenerationImage-to-Image Translation

OracleAnalyser: Analysing Implicit Semantics of Oracle Bone Scripts through MLLMs with Post-training

2026-06-24 · Zijia Song, Yelin Wang, Zhengyi Ma, Zitong Yu 외 arxiv

With the advancement of artificial intelligence, research on oracle bone scripts has entered a new era. However, existing methods and benchmarks remain largely confined to recognition tasks, overlooking the equally cruci…

Oracle Bone Inscriptions Multi-modal Dataset

2024-07-04 · Bang Li, Donghao Luo, Yujie Liang, Jing Yang 외

Oracle bone inscriptions(OBI) is the earliest developed writing system in China, bearing invaluable written exemplifications of early Shang history and paleography. However, the task of deciphering OBI, in the current cl…

DeciphermentDenoising