paper-with-me

홈 › Papers

Ocean-OCR: Towards General OCR Application via a Vision-Language Model

2025-01-26 · Song Chen, Xinyu Guo, Yadong Li, Tao Zhang, MingAn Lin, Dongdong Kuang, Youwei Zhang, Lingfeng Ming, Fengyu Zhang, Yuran Wang, Jianhua Xu, Zenan Zhou, WeiPeng Chen

Multimodal large language models (MLLMs) have shown impressive capabilities across various domains, excelling in processing and understanding information from multiple modalities. Despite the rapid progress made previously, insufficient OCR ability hinders MLLMs from excelling in text-related tasks. In this paper, we present \textbf{Ocean-OCR}, a 3B MLLM with state-of-the-art performance on various OCR scenarios and comparable understanding ability on general tasks. We employ Native Resolution ViT to enable variable resolution input and utilize a substantial collection of high-quality OCR datasets to enhance the model performance. We demonstrate the superiority of Ocean-OCR through comprehensive experiments on open-source OCR benchmarks and across various OCR scenarios. These scenarios encompass document understanding, scene text recognition, and handwritten recognition, highlighting the robust OCR capabilities of Ocean-OCR. Note that Ocean-OCR is the first MLLM to outperform professional OCR models such as TextIn and PaddleOCR.

📄 PDF Abstract BibTeX arXiv:2501.15558

Code (1)

guoxy25/Ocean-OCR 공식 구현 pytorch

Tasks

document understandingLanguage ModelingLanguage ModellingOptical Character Recognition (OCR)Scene Text Recognition

Similar Papers 제목 키워드 기반

Composing Open-domain Vision with RAG for Ocean Monitoring and Conservation

2024-12-03 · Sepand Dyanatkar, Angran Li, Alexander Dungate

Climate change's destruction of marine biodiversity is threatening communities and economies around the world which rely on healthy oceans for their livelihoods. The challenge of applying computer vision to niche, real-w…

RAGRetrievalRetrieval-augmented Generation

OceanAI: A Conversational Platform for Accurate, Transparent, Near-Real-Time Oceanographic Insights

2025-11-02 · Bowen Chen, Jayesh Gajbhar, Gregory Dusek, Rob Redmon 외 arxiv

Artificial intelligence is transforming the sciences, yet general conversational AI systems often generate unverified "hallucinations" undermining scientific rigor. We present OceanAI, a conversational platform that inte…

MarineGPT: Unlocking Secrets of Ocean to the Public

2023-10-20 · Ziqiang Zheng, Jipeng Zhang, Tuan-Anh Vu, Shizhe Diao 외

Large language models (LLMs), such as ChatGPT/GPT-4, have proven to be powerful tools in promoting the user experience as an AI assistant. The continuous works are proposing multi-modal large language models (MLLM), empo…

Language Modelling

Optical Ocean Recipes: Creating Realistic Datasets to Facilitate Underwater Vision Research

2025-09-24 · Patricia Schöntag, David Nakath, Judith Fischer, Rüdiger Röttgers 외 arxiv

The development and evaluation of machine vision in underwater environments remains challenging, often relying on trial-and-error-based testing tailored to specific applications. This is partly due to the lack of control…

Image Restoration

OceanPile: A Large-Scale Multimodal Ocean Corpus for Foundation Models

2026-04-25 · Yida Xue, Ningyu Zhang, Tingwei Wu, Zhe Ma 외 arxiv

The vast and underexplored ocean plays a critical role in regulating global climate and supporting marine biodiversity, yet artificial intelligence has so far delivered limited impact in this domain due to a fundamental …