paper-with-me

Papers

ZeroSense:How Vision matters in Long Context Compression

2026-03-12 · Yonghan Gao, Zehong Chen, Lijian Xu, Jingzhi Chen, Jingwei Guan, Xingyu Zeng arxiv

Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image rendering. However, existing evaluation protocols heavily rely on downstream task performance. Such evaluation metrics fail to accurately measure text preservation due to the strong inherent linguistic priors of Multimodal Large Language Models (MLLMs). In this work, we introduce a new evaluation framework that decouples MLLMs' capabilities to faithfully assess VTC quality. Within this framework, we further introduce the ZeroSense Benchmark to ensure low semantic correlation of testing samples. By eliminating contextual dependencies, our benchmark guarantees that the evaluation results are purely reflective of VTC quality, unaffected by the semantic inference capabilities of downstream models. Extensive experiments across multiple datasets demonstrate that VTC quality and downstream task accuracy diverge significantly, highlighting the necessity of our decoupled evaluation framework.

📄 PDF Abstract BibTeX arXiv:2603.11846

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

Data Distribution Matters: A Data-Centric Perspective on Context Compression for Large Language Model

2026-02-02 · Kangtao Lv, Jiwei Tang, Langming Liu, Haibin Chen 외 arxiv

The deployment of Large Language Models (LLMs) in long-context scenarios is hindered by computational inefficiency and significant information redundancy. Although recent advancements have widely adopted context compress…

Scene Matters: Model-based Deep Video Compression

2023-03-08 · ICCV 2023 1 · Lv Tang, Xinfeng Zhang, Gai Zhang, Xiaoqi Ma

Video compression has always been a popular research area, where many traditional and deep video compression methods have been proposed. These methods typically rely on signal prediction theory to enhance compression per…

modelVideo Compression

Can Vision-Language Models Handle Long-Context Code? An Empirical Study on Visual Compression

2026-01-31 · Jianping Zhong, Guochang Li, Chen Zhi, Junxiao Han 외 arxiv

Large Language Models (LLMs) struggle with long-context code due to window limitations. Existing textual code compression methods mitigate this via selective filtering but often disrupt dependency closure, causing semant…

Question AnsweringCode Completion

Optical Context Compression Is Just (Bad) Autoencoding

2025-12-03 · Ivan Yee Lee, Cheng Yang, Taylor Berg-Kirkpatrick arxiv

DeepSeek-OCR shows that rendered text can be reconstructed from a small number of vision tokens, sparking excitement about using vision as a compression medium for long textual contexts. But this pipeline requires render…

Locality Matters for Training-Free Audio Token Compression in Audio-Language Models

2026-05-24 · Jiale Luo, Xiaoyu Liang, Haoji Hu arxiv

Audio-language models (ALMs) are increasingly used for audio captioning, question answering, and open-ended audio understanding, but their inference cost remains high when audio inputs are represented as long prefix-toke…

Question AnsweringAudio captioning