paper-with-me

홈 › Papers

UniLipi: A Unified Multi-Script OCR for Historical Indic Manuscripts

2026-08-28 · Tathagata Ghosh, Sai Madhusudan Gunda, Simran Singh Sandral, Ravi Kiran Sarvadevabhatla arxiv

Optical character recognition (OCR) for handwritten Indic manuscripts is essential for large-scale digitization and computational access to manuscript heritage. However, existing approaches are typically developed for one script at a time and require substantial script-specific customization. This limits scalability and practical deployment across diverse collections. We present UniLipi, a unified multi-script OCR model for handwritten Indic manuscripts trained jointly across 13 Indic scripts within a single framework. UniLipi directly handles realistic manuscript conditions, including extreme variation in line geometry, large variation in line length, and partial interruptions caused by non-textual manuscript entities such as holes, stains, or pictorial illustrations. To operate effectively under ultra low-resource conditions, the model leverages script-aware synthetic manuscript data generation, substantially reducing reliance on large volumes of real annotated data. Beyond historical manuscripts, we show that UniLipi serves as an effective foundational pretrained model. Specifically, its learned representations enable good OCR performance for contemporary Indic handwriting and extend to several non-Indic scripts, including Tibetan, Italian, Latin, and Chinese scripts. In addition to transcription, UniLipi predicts script identity and per-line native character counts, supporting practical manuscript cataloging workflows.

📄 PDF Abstract BibTeX arXiv:2608.28195

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Indiscapes: Instance Segmentation Networks for Layout Parsing of Historical Indic Manuscripts

2019-12-15 · Abhishek Prusty, Sowmya Aitha, Abhishek Trivedi, Ravi Kiran Sarvadevabhatla

Historical palm-leaf manuscript and early paper documents from Indian subcontinent form an important part of the world's literary and cultural heritage. Despite their importance, large-scale annotated Indic manuscript im…

DiversityInstance SegmentationOptical Character Recognition (OCR)Semantic Segmentation

Unified NMT models for the Indian subcontinent transcending script-barriers

2021-11-16 · ACL ARR November 2021 11 · Anonymous

Highly accurate machine translation systems are very important in societies and countries where multilinguality is very common, and where English often does not suffice. The Indian subcontinent is such a region, with all…

Machine TranslationNMTTranslation

The Effects of Character-Level Data Augmentation on Style-Based Dating of Historical Manuscripts

2022-12-15 · Lisa Koopmans, Maruf A. Dhali, Lambert Schomaker

Identifying the production dates of historical manuscripts is one of the main goals for paleographers when studying ancient documents. Automatized methods can provide paleographers with objective tools to estimate dates …

Data Augmentation

InteChar: A Unified Oracle Bone Character List for Ancient Chinese Language Modeling

2025-08-12 · Xiaolei Diao, Zhihan Zhou, Lida Shi, Ting Wang 외 arxiv

Constructing historical language models (LMs) plays a crucial role in aiding archaeological provenance studies and understanding ancient cultures. However, existing resources present major challenges for training effecti…

Data Augmentation

Unveiling Text in Challenging Stone Inscriptions: A Character-Context-Aware Patching Strategy for Binarization

2026-01-07 · Pratyush Jena, Amal Joseph, Arnav Sharma, Ravi Kiran Sarvadevabhatla arxiv

Binarization is a popular first step towards text extraction in historical artifacts. Stone inscription images pose severe challenges for binarization due to poor contrast between etched characters and the stone backgrou…

Zero-shot Generalization