paper-with-me

Optical Character Recognition (OCR)

6개 벤치마크 · 논문 1,209편 · 이 태스크의 논문 보기 →

Benchmarks

FSNS - Test

결과 3개

I2L-140K

결과 2개

SUT

결과 2개

im2latex-100k

결과 1개

Most implemented

Papers

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning

2025-07-17 · Senqiao Yang, Junyi Li, Xin Lai, Bei Yu 외

Recent advancements in vision-language models (VLMs) have improved performance by increasing the number of visual tokens, which are often significantly longer than text tokens. However, we observe that most real-world sc…

Language ModelingLanguage ModellingOptical Character Recognition (OCR)reinforcement-learning+2

DeQA-Doc: Adapting DeQA-Score to Document Image Quality Assessment

2025-07-17 · Junjie Gao, Runze Liu, Yingzhe Peng, Shujian Yang 외

Document quality assessment is critical for a wide range of applications including document digitization, OCR, and archival. However, existing approaches often struggle to provide accurate and robust quality scores, limi…

Document Image Quality AssessmentImage Quality AssessmentOptical Character Recognition (OCR)

Seeing the Signs: A Survey of Edge-Deployable OCR Models for Billboard Visibility Analysis

2025-07-15 · Maciej Szankin, Vidhyananth Venkatasamy, Lihang Ying

Outdoor advertisements remain a critical medium for modern marketing, yet accurately verifying billboard text visibility under real-world conditions is still challenging. Traditional Optical Character Recognition (OCR) p…

MarketingOptical Character RecognitionOptical Character Recognition (OCR)Scene Understanding

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends

2025-07-14 · Yihao Ding, Siwen Luo, Yue Dai, Yanbei Jiang 외

Visually-Rich Document Understanding (VRDU) has emerged as a critical field, driven by the need to automatically process documents containing complex visual, textual, and layout information. Recently, Multimodal Large La…

document understandingOptical Character RecognitionOptical Character Recognition (OCR)

Design and Implementation of an OCR-Powered Pipeline for Table Extraction from Invoices

2025-07-09 · Parshva Dhilankumar Patel

This paper presents the design and development of an OCR-powered pipeline for efficient table extraction from invoices. The system leverages Tesseract OCR for text recognition and custom post-processing logic to detect, …

Boundary DetectionOptical Character Recognition (OCR)Table Extraction

Orchestrator-Agent Trust: A Modular Agentic AI Visual Classification System with Trust-Aware Orchestration and RAG-Based Reasoning

2025-07-09 · Konstantinos I. Roumeliotis, Ranjan Sapkota, Manoj Karkee, Nikolaos D. Tselikas

Modern Artificial Intelligence (AI) increasingly relies on multi-agent architectures that blend visual and language understanding. Yet, a pressing challenge remains: How can we trust these agents especially in zero-shot …

BenchmarkingImage RetrievalOptical Character Recognition (OCR)RAG+3

전체 1,209편 보기 →