Optical Character Recognition (OCR)
6개 벤치마크 · 논문 1,209편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
An End-to-End Trainable Neural Network for Image-based Sequence Recognition and Its Application to Scene Text Recognition
EAST: An Efficient and Accurate Scene Text Detector
Shape Robust Text Detection with Progressive Scale Expansion Network
Real-time Scene Text Detection with Differentiable Binarization
Image-to-Markup Generation with Coarse-to-Fine Attention
PP-OCR: A Practical Ultra Lightweight OCR System
Papers
VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning
Recent advancements in vision-language models (VLMs) have improved performance by increasing the number of visual tokens, which are often significantly longer than text tokens. However, we observe that most real-world sc…
Language ModelingLanguage ModellingOptical Character Recognition (OCR)reinforcement-learning+2DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment
Document quality assessment is critical for a wide range of applications including document digitization, OCR, and archival. However, existing approaches often struggle to provide accurate and robust quality scores, limi…
Document Image Quality AssessmentImage Quality AssessmentOptical Character Recognition (OCR)Seeing the Signs: A Survey of Edge-Deployable OCR Models for Billboard Visibility Analysis
Outdoor advertisements remain a critical medium for modern marketing, yet accurately verifying billboard text visibility under real-world conditions is still challenging. Traditional Optical Character Recognition (OCR) p…
MarketingOptical Character RecognitionOptical Character Recognition (OCR)Scene UnderstandingA Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends
Visually-Rich Document Understanding (VRDU) has emerged as a critical field, driven by the need to automatically process documents containing complex visual, textual, and layout information. Recently, Multimodal Large La…
document understandingOptical Character RecognitionOptical Character Recognition (OCR)Design and Implementation of an OCR-Powered Pipeline for Table Extraction from Invoices
This paper presents the design and development of an OCR-powered pipeline for efficient table extraction from invoices. The system leverages Tesseract OCR for text recognition and custom post-processing logic to detect, …
Boundary DetectionOptical Character Recognition (OCR)Table ExtractionOrchestrator-Agent Trust: A Modular Agentic AI Visual Classification System with Trust-Aware Orchestration and RAG-Based Reasoning
Modern Artificial Intelligence (AI) increasingly relies on multi-agent architectures that blend visual and language understanding. Yet, a pressing challenge remains: How can we trust these agents especially in zero-shot …
BenchmarkingImage RetrievalOptical Character Recognition (OCR)RAG+3