paper-with-me

Papers

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models

2025-09-07 · Ruiqi Shen, Haotian Wu, Wenjing Zhang, Jiangjing Hu, Deniz Gunduz arxiv

Recent deep learning-based methods for lossy image compression achieve competitive rate-distortion performance through extensive end-to-end training and advanced architectures. However, emerging applications increasingly prioritize semantic preservation over pixel-level reconstruction and demand robust performance across diverse data distributions and downstream tasks. These challenges call for advanced semantic compression paradigms. Motivated by the zero-shot and representational capabilities of multimodal foundation models, we propose a novel semantic compression method based on the contrastive language-image pretraining (CLIP) model. Rather than compressing images for reconstruction, we propose compressing the CLIP feature embeddings into minimal bits while preserving semantic information across different tasks. Experiments show that our method maintains semantic integrity across benchmark datasets, achieving an average bit rate of approximately 2-3* 10(-3) bits per pixel. This is less than 5% of the bitrate required by mainstream image compression approaches for comparable performance. Remarkably, even under extreme compression, the proposed approach exhibits zero-shot robustness across diverse data distributions and downstream tasks.

📄 PDF Abstract BibTeX arXiv:2509.05925

Code (0)

등록된 구현이 없습니다.

Tasks

Image Compression

Similar Papers 제목 키워드 기반

A Spatial RNN Codec for End-to-End Image Compression

2020-06-01 · CVPR 2020 6 · Chaoyi Lin, Jiabao Yao, Fangdong Chen, Li Wang

Recently, deep learning has been explored as a promising direction for image compression. Removing the spatial redundancy of the image is crucial for image compression and most learning based methods focus on removing th…

Image CompressionMS-SSIMSSIM

Semantic Compression via Multimodal Representation Learning

2025-09-29 · Eleonora Grassucci, Giordano Cicchetti, Aurelio Uncini, Danilo Comminiello arxiv

Multimodal representation learning produces high-dimensional embeddings that align diverse modalities in a shared latent space. While this enables strong generalization, it also introduces scalability challenges, both in…

Representation Learning

Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

2024-09-19 · Zuyan Liu, Yuhao Dong, Ziwei Liu, Winston Hu 외

Visual data comes in various forms, ranging from small icons of just a few pixels to long videos spanning hours. Existing multi-modal LLMs usually standardize these diverse visual inputs to a fixed resolution for visual …

document understandingVideo Question Answering

Attention-guided Image Compression by Deep Reconstruction of Compressive Sensed Saliency Skeleton

2021-03-29 · CVPR 2021 1 · Xi Zhang, Xiaolin Wu

We propose a deep learning system for attention-guided dual-layer image compression (AGDL). In the AGDL compression system, an image is encoded into two layers, a base layer and an attention-guided refinement layer. Unli…

Compressive SensingImage Compression

Rule of Three for Superresolution of Still Images with Applications to Compression and Denoising

2014-05-04 · Mario Mastriani

We describe a new method for superresolution of still images (in the wavelet domain) based on the reconstruction of missing details subbands pixels at a given ith level via Rule of Three (Ro3) between pixels of approxima…

Denoising