paper-with-me

홈 › Papers

RTPrune: Reading-Twice Inspired Token Pruning for Efficient DeepSeek-OCR Inference

2026-05-01 · Ben Wan, Yan Feng, Zihan Tang, Weizhe Huang, Yuting Zeng, Jia Wang, Tongxuan Liu arxiv

DeepSeek-OCR leverages visual-text compression to reduce long-text processing costs and accelerate inference, yet visual tokens remain prone to redundant textual and structural information. Moreover, current token pruning methods for conventional vision-language models (VLMs) fail to preserve textual fidelity due to improper compression mechanisms. By analyzing the decoding process of DeepSeek-OCR, we find that a distinct two-stage reading trajectory: the model initially prioritizes the majority of high-norm tokens, then subsequently redistributes its attention to the remaining ones. Motivated by this insight, we propose RTPrune, a two-stage token pruning method tailored for DeepSeek-OCR. In the first stage, we prioritize high-norm visual tokens that capture salient textual and structural information. In the second stage, the remaining tokens are paired and merged based on optimal transport theory to achieve efficient feature aggregation. We further introduce a dynamic pruning ratio that adapts to token similarity and textual density for OCR tasks, enabling a better efficiency-accuracy trade-off. Extensive experiments demonstrate state-of-the-art performance, as evidenced by 99.47% accuracy and 1.23$\times$ faster prefill on OmniDocBench, achieved with 84.25% token retention when applied to DeepSeek-OCR-Large.

📄 PDF Abstract BibTeX arXiv:2605.00392

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

A Glimpse to Compress: Dynamic Visual Token Pruning for Large Vision-Language Models

2025-08-03 · Quan-Sheng Zeng, Yunheng Li, Qilong Wang, Peng-Tao Jiang 외 arxiv

Visual token compression is critical for Large Vision-Language Models (LVLMs) to efficiently process high-resolution inputs. Existing methods that typically adopt fixed compression ratios cannot adapt to scenes of varyin…

Answer Generation

Nüwa: Mending the Spatial Integrity Torn by VLM Token Pruning

2026-02-03 · Yihong Huang, Fei Ma, Yihua Shao, Jingcai Guo 외 arxiv

Vision token pruning has proven to be an effective acceleration technique for the efficient Vision Language Model (VLM). However, existing pruning methods demonstrate excellent performance preservation in visual question…

Visual Question AnsweringSemantic SimilarityVisual Grounding

Bridging The Gaps Between Token Pruning and Full Pre-training via Masked Fine-tuning

2023-10-26 · Fengyuan Shi, LiMin Wang

Despite the success of transformers on various computer vision tasks, they suffer from excessive memory and computational cost. Some works present dynamic vision transformers to accelerate inference by pruning redundant …

Pruning the Index Contents for Memory Efficient Open-Domain QA

2021-02-21 · Martin Fajcik, Martin Docekal, Karel Ondrej, Pavel Smrz

This work presents a novel pipeline that demonstrates what is achievable with a combined effort of state-of-the-art approaches. Specifically, it proposes the novel R2-D2 (Rank twice, reaD twice) pipeline composed of retr…

Open-Domain Question Answering

Exploring Token Pruning in Vision State Space Models

2024-09-27 · Zheng Zhan, Zhenglun Kong, Yifan Gong, Yushu Wu 외

State Space Models (SSMs) have the advantage of keeping linear computational complexity compared to attention modules in transformers, and have been applied to vision tasks as a new type of powerful vision foundation mod…

State Space Models