paper-with-me

홈 › Papers

KOSMOS-2.5: A Multimodal Literate Model

2023-09-20 · Tengchao Lv, Yupan Huang, Jingye Chen, Yuzhong Zhao, Yilin Jia, Lei Cui, Shuming Ma, Yaoyao Chang, Shaohan Huang, Wenhui Wang, Li Dong, Weiyao Luo, Shaoxiang Wu, Guoxin Wang, Cha Zhang, Furu Wei

The automatic reading of text-intensive images represents a significant advancement toward achieving Artificial General Intelligence (AGI). In this paper we present KOSMOS-2.5, a multimodal literate model for machine reading of text-intensive images. Pre-trained on a large-scale corpus of text-intensive images, KOSMOS-2.5 excels in two distinct yet complementary transcription tasks: (1) generating spatially-aware text blocks, where each block of text is assigned spatial coordinates within the image, and (2) producing structured text output that captures both style and structure in markdown format. This unified multimodal literate capability is achieved through a shared decoder-only autoregressive Transformer architecture and task-specific prompts. Building on this foundation, we fine-tune KOSMOS-2.5 for document understanding tasks, resulting in a document understanding generalist named KOSMOS-2.5-CHAT. Additionally, a large corpus of 357.4 million document pages spanning diverse domains was curated for pre-training. We evaluate KOSMOS-2.5 on two newly proposed benchmarks, OCREval and MarkdownEval, for document-level text recognition and image-to-markdown generation, demonstrating impressive literate capabilities comparable to GPT-4o. KOSMOS-2.5-CHAT achieves performance comparable to other state-of-the-art generalists that are five times larger (1.3B vs. 7B) across nine text-rich visual question answering benchmarks. Models and code have been available at \url{https://aka.ms/kosmos25}.

📄 PDF Abstract BibTeX arXiv:2309.11419

Code (0)

등록된 구현이 없습니다.

Tasks

document understandingmodelQuestion AnsweringReading ComprehensionText GenerationVisual Question Answering

Methods 이 논문이 사용한 방법론

Multi-Head Attention 설명 없음
Attention 설명 없음
Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Dense Connections Dense Connections, or Fully Connected Connections, are a type of layer in a deep neural network that use a linear operation where every input is connected to every output…
Position-Wise Feed-Forward Layer 설명 없음
Absolute Position Encodings Absolute Position Encodings are a type of position embeddings for [Transformer-based models] where positional encodings are…
Residual Connection 설명 없음
Adam 설명 없음

Similar Papers 제목 키워드 기반

Kosmos-2: Grounding Multimodal Large Language Models to the World

2023-06-26 · Zhiliang Peng, Wenhui Wang, Li Dong, Yaru Hao 외

We introduce Kosmos-2, a Multimodal Large Language Model (MLLM), enabling new capabilities of perceiving object descriptions (e.g., bounding boxes) and grounding text to the visual world. Specifically, we represent refer…

Image CaptioningIn-Context LearningLanguage ModelingLanguage Modelling+9

Kosmos-G: Generating Images in Context with Multimodal Large Language Models

2023-10-04 · Xichen Pan, Li Dong, Shaohan Huang, Zhiliang Peng 외

Recent advancements in subject-driven image generation have made significant strides. However, current methods still fall short in diverse application scenarios, as they require test-time tuning and cannot accept interle…

DecoderImage Generation

Language Is Not All You Need: Aligning Perception with Language Models

2023-02-27 · NeurIPS 2023 11 · Shaohan Huang, Li Dong, Wenhui Wang, Yaru Hao 외

A big convergence of language, multimodal perception, action, and world modeling is a key step toward artificial general intelligence. In this work, we introduce Kosmos-1, a Multimodal Large Language Model (MLLM) that ca…

AllImage CaptioningLanguage ModelingLanguage Modelling+5

Kosmos: An AI Scientist for Autonomous Discovery

2025-11-04 · Ludovico Mitchener, Angela Yiu, Benjamin Chang, Mathieu Bourdenx 외 arxiv

Data-driven scientific discovery requires iterative cycles of literature search, hypothesis generation, and data analysis. Substantial progress has been made towards AI agents that can automate scientific research, but a…

KOSMOS: Knowledge-graph Oriented Social media and Mainstream media Overview System

2020-12-11 · Chua Hao Yang, Yong Shan Jie, Boon Kok Chin, Lander Chin 외

We introduce KOSMOS, a knowledge retrieval system based on the constructed knowledge graph of social media and mainstream media documents. The system first identifies key events from the documents at each time frame thro…

ArticlesClusteringEntity DisambiguationRetrieval