paper-with-me

Papers

Compression Tells Intelligence: Visual Coding, Visual Token Technology, and the Unification

2026-01-28 · Xin Jin, Jinming Liu, Yuntao Wei, Junyan Lin, Zhicheng Wang, Jianguo Huang, Xudong Yang, Yanxiao Liu, Wenjun Zeng arxiv

"Compression Tells Intelligence", is supported by research in artificial intelligence, particularly concerning (multimodal) large language models (LLMs/MLLMs), where compression efficiency often correlates with improved model performance and capabilities. For compression, classical visual coding based on traditional information theory has developed over decades, achieving great success with numerous international industrial standards widely applied in multimedia (e.g., image/video) systems. Except that, the recent emergingvisual token technology of generative multi-modal large models also shares a similar fundamental objective like visual coding: maximizing semantic information fidelity during the representation learning while minimizing computational cost. Therefore, this paper provides a comprehensive overview of two dominant technique families first -- Visual Coding and Vision Token Technology -- then we further unify them from the aspect of optimization, discussing the essence of compression efficiency and model performance trade-off behind. Next, based on the proposed unified formulation bridging visual coding andvisual token technology, we synthesize bidirectional insights of themselves and forecast the next-gen visual codec and token techniques. Last but not least, we experimentally show a large potential of the task-oriented token developments in the more practical tasks like multimodal LLMs (MLLMs), AI-generated content (AIGC), and embodied AI, as well as shedding light on the future possibility of standardizing a general token technology like the traditional codecs (e.g., H.264/265) with high efficiency for a wide range of intelligent tasks in a unified and effective manner.

📄 PDF Abstract BibTeX arXiv:2601.20742

Code (0)

등록된 구현이 없습니다.

Tasks

Representation Learning

Similar Papers 제목 키워드 기반

Emerging Advances in Learned Video Compression: Models, Systems and Beyond

2025-04-30 · Chuanmin Jia, Feng Ye, Siwei Ma, Wen Gao 외

Video compression is a fundamental topic in the visual intelligence, bridging visual signal sensing/capturing and high-level visual analytics. The broad success of artificial intelligence (AI) technology has enriched the…

Video Compression

Beyond GFVC: A Progressive Face Video Compression Framework with Adaptive Visual Tokens

2024-10-11 · Bolin Chen, Shanzhi Yin, Zihan Zhang, Jie Chen 외

Recently, deep generative models have greatly advanced the progress of face video coding towards promising rate-distortion performance and diverse application functionalities. Beyond traditional hybrid video coding parad…

Motion EstimationPhilosophyVideo Compression

Scalable Face Image Coding via StyleGAN Prior: Towards Compression for Human-Machine Collaborative Vision

2023-12-25 · Qi Mao, Chongyu Wang, Meng Wang, Shiqi Wang 외

The accelerated proliferation of visual content and the rapid development of machine vision technologies bring significant challenges in delivering visual data on a gigantic scale, which shall be effectively represented …

Image Compression

LLaVA-UHD v4: What Makes Efficient Visual Encoding in MLLMs?

2026-05-09 · Kechen Fang, Yihua Qin, Chongyi Wang, Wenshuo Ma 외 arxiv

Visual encoding constitutes a major computational bottleneck in Multimodal Large Language Models (MLLMs), especially for high-resolution image inputs. The prevailing practice typically adopts global encoding followed by …

Towards Analysis-friendly Face Representation with Scalable Feature and Texture Compression

2020-04-21 · Shurun Wang, Shiqi Wang, Wenhan Yang, Xinfeng Zhang 외

It plays a fundamental role to compactly represent the visual information towards the optimization of the ultimate utility in myriad visual data centered applications. With numerous approaches proposed to efficiently com…

Image Compression