LMM-driven Semantic Image-Text Coding for Ultra Low-bitrate Learned Image Compression
Supported by powerful generative models, low-bitrate learned image compression (LIC) models utilizing perceptual metrics have become feasible. Some of the most advanced models achieve high compression rates and superior perceptual quality by using image captions as sub-information. This paper demonstrates that using a large multi-modal model (LMM), it is possible to generate captions and compress them within a single model. We also propose a novel semantic-perceptual-oriented fine-tuning method applicable to any LIC network, resulting in a 41.58\% improvement in LPIPS BD-rate compared to existing methods. Our implementation and pre-trained weights are available at https://github.com/tokkiwa/ImageTextCoding.
Code (1)
Tasks
Image CaptioningImage CompressionSimilar Papers 제목 키워드 기반
Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
We consider the problem of ultra-low bit rate visual communication for remote vision analysis, human interactions and control in challenging scenarios with very low communication bandwidth, such as deep space exploration…
Text-to-Image GenerationImage ReconstructionImage CompressionRobot NavigationTaming Large Multimodal Agents for Ultra-low Bitrate Semantically Disentangled Image Compression
It remains a significant challenge to compress images at ultra-low bitrate while achieving both semantic consistency and high perceptual quality. We propose a novel image compression framework, Semantically Disentangled …
DecoderImage CompressionObjectGenerative Latent Coding for Ultra-Low Bitrate Image and Video Compression
Most existing approaches for image and video compression perform transform coding in the pixel space to reduce redundancy. However, due to the misalignment between the pixel-space distortion and human perception, such sc…
Image CompressionVideo CompressionYour Demands Deserve More Bits: Referring Semantic Image Compression at Ultra-low Bitrate
With the help of powerful generative models, Semantic Image Compression (SIC) has achieved impressive performance at ultra-low bitrate. However, due to coarse-grained visual-semantic alignment and inherent randomness, th…
Image CompressionLow-Rate Semantic Communication with Codebook-based Conditional Generative Models
Generative semantic communication models are reshaping semantic communication frameworks by moving beyond pixel-wise optimization to align with human perception. However, many existing approaches prioritize image-level p…
Saliency DetectionSemantic Communication