paper-with-me

Papers

PsOCR: Benchmarking Large Multimodal Models for Optical Character Recognition in Low-resource Pashto Language

2025-05-15 · Ijazul Haq, Yingjie Zhang, Irfan Ali Khan

This paper evaluates the performance of Large Multimodal Models (LMMs) on Optical Character Recognition (OCR) in the low-resource Pashto language. Natural Language Processing (NLP) in Pashto faces several challenges due to the cursive nature of its script and a scarcity of structured datasets. To address this, we developed a synthetic Pashto OCR dataset, PsOCR, consisting of one million images annotated with bounding boxes at word, line, and document levels, suitable for training and evaluating models based on different architectures, including Convolutional Neural Networks (CNNs) and Transformers. PsOCR covers variations across 1,000 unique font families, colors, image sizes, and layouts. A benchmark subset of 10K images was selected to evaluate the performance of several LMMs, including seven open-source models: DeepSeek's Janus, InternVL, MiniCPM, Florence, and Qwen (3B and 7B), and four closed-source models: GPT-4o, Gemini, Claude, and Grok. Experimental results demonstrate that Gemini achieves the best performance among all models, whereas among open-source models, Qwen-7B stands out. This work provides an insightful assessment of the capabilities and limitations of current LMMs for OCR tasks in Pashto and establishes a foundation for further research not only in Pashto OCR but also for other similar scripts such as Arabic, Persian, and Urdu. PsOCR is available at https://github.com/zirak-ai/PashtoOCR.

📄 PDF Abstract BibTeX arXiv:2505.10055

Code (1)

zirak-ai/pashtoocr 공식 구현

Tasks

BenchmarkingOptical Character RecognitionOptical Character Recognition (OCR)

Methods 이 논문이 사용한 방법론

Florence Florence is a computer vision foundation model aiming to learn universal visual-language representations that be adapted to various computer vision tasks, visual question…

Similar Papers 제목 키워드 기반

IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs

2025-11-06 · Ali Faraz, Akash, Shaharukh Khan, Raja Kolla 외 arxiv

Vision-language models (VLMs) have demonstrated impressive generalization across multimodal tasks, yet most evaluation benchmarks remain Western-centric, leaving open questions about their performance in culturally diver…

Multimodal Machine TranslationVisual Question Answering

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

2024-12-31 · Ling Fu, Biao Yang, Zhebin Kuang, Jiajun Song 외

Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest recently. Existing benchmarks have highlighted the impressive performance of LMMs in text reco…

BenchmarkingLogical ReasoningOptical Character RecognitionOptical Character Recognition (OCR)+1

Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments

2025-02-10 · Sankalp Nagaonkar, Augustya Sharma, Ashish Choithani, Ashutosh Trivedi

This paper introduces an open-source benchmark for evaluating Vision-Language Models (VLMs) on Optical Character Recognition (OCR) tasks in dynamic video environments. We present a curated dataset containing 1,477 manual…

BenchmarkingOptical Character RecognitionOptical Character Recognition (OCR)

CC-OCR V2: Benchmarking Large Multimodal Models for Literacy in Real-world Document Processing

2026-05-05 · Zhipeng Xu, Junhao Ji, Zulong Chen, Zhenghao Liu 외 arxiv

Large Multimodal Models (LMMs) have recently shown strong performance on Optical Character Recognition (OCR) tasks, demonstrating their promising capability in document literacy. However, their effectiveness in real-worl…

Key Information ExtractionQuestion Answering

Benchmarking the Robustness of Optical Flow Estimation to Corruptions

2024-11-22 · Zhonghua Yi, Hao Shi, Qi Jiang, Yao Gao 외

Optical flow estimation is extensively used in autonomous driving and video editing. While existing models demonstrate state-of-the-art performance across various benchmarks, the robustness of these methods has been infr…

Autonomous DrivingBenchmarkingOptical Flow EstimationVideo Editing