Deciphering the Underserved: Benchmarking LLM OCR for Low-Resource Scripts
This study investigates the potential of Large Language Models (LLMs), particularly GPT-4o, for Optical Character Recognition (OCR) in low-resource scripts such as Urdu, Albanian, and Tajik, with English serving as a benchmark. Using a meticulously curated dataset of 2,520 images incorporating controlled variations in text length, font size, background color, and blur, the research simulates diverse real-world challenges. Results emphasize the limitations of zero-shot LLM-based OCR, particularly for linguistically complex scripts, highlighting the need for annotated datasets and fine-tuned models. This work underscores the urgency of addressing accessibility gaps in text digitization, paving the way for inclusive and robust OCR solutions for underserved languages.
Code (1)
Tasks
BenchmarkingOptical Character RecognitionOptical Character Recognition (OCR)Similar Papers 제목 키워드 기반
Diff-Oracle: Deciphering Oracle Bone Scripts with Controllable Diffusion Model
Deciphering oracle bone scripts plays an important role in Chinese archaeology and philology. However, a significant challenge remains due to the scarcity of oracle character images. To overcome this issue, we propose Di…
Image GenerationImage-to-Image TranslationReasoning Over the Glyphs: Evaluation of LLM's Decipherment of Rare Scripts
We explore the capabilities of LVLMs and LLMs in deciphering rare scripts not encoded in Unicode. We introduce a novel approach to construct a multimodal dataset of linguistic puzzles involving such scripts, utilizing a …
DeciphermentAsymmetrical Reciprocity-based Federated Learning for Resolving Disparities in Medical Diagnosis
Geographic health disparities pose a pressing global challenge, particularly in underserved regions of low- and middle-income nations. Addressing this issue requires a collaborative approach to enhance healthcare quality…
DiagnosticFederated Learningimage-classificationImage Classification+3A Contrastive Pre-trained Foundation Model for Deciphering Imaging Noisomics across Modalities
Characterizing imaging noise is notoriously data-intensive and device-dependent, as modern sensors entangle physical signals with complex algorithmic artifacts. Current paradigms struggle to disentangle these factors wit…
Zero-shot GeneralizationContrastive LearningDistinguishing Repetition Disfluency from Morphological Reduplication in Bangla ASR Transcripts: A Novel Corpus and Benchmarking Analysis
Automatic Speech Recognition (ASR) transcripts, especially in low-resource languages like Bangla, contain a critical ambiguity: word-word repetitions can be either Repetition Disfluency (unintentional ASR error/hesitatio…
Speech Recognition