paper-with-me

Papers

Cross Modal Compression: Towards Human-comprehensible Semantic Compression

2022-09-06 · Jiguo Li, Chuanmin Jia, Xinfeng Zhang, Siwei Ma, Wen Gao

Traditional image/video compression aims to reduce the transmission/storage cost with signal fidelity as high as possible. However, with the increasing demand for machine analysis and semantic monitoring in recent years, semantic fidelity rather than signal fidelity is becoming another emerging concern in image/video compression. With the recent advances in cross modal translation and generation, in this paper, we propose the cross modal compression~(CMC), a semantic compression framework for visual data, to transform the high redundant visual data~(such as image, video, etc.) into a compact, human-comprehensible domain~(such as text, sketch, semantic map, attributions, etc.), while preserving the semantic. Specifically, we first formulate the CMC problem as a rate-distortion optimization problem. Secondly, we investigate the relationship with the traditional image/video compression and the recent feature compression frameworks, showing the difference between our CMC and these prior frameworks. Then we propose a novel paradigm for CMC to demonstrate its effectiveness. The qualitative and quantitative results show that our proposed CMC can achieve encouraging reconstructed results with an ultrahigh compression ratio, showing better compression performance than the widely used JPEG baseline.

📄 PDF Abstract BibTeX arXiv:2209.02574

Code (0)

등록된 구현이 없습니다.

Tasks

Feature CompressionSemantic CompressionVideo Compression

Similar Papers 제목 키워드 기반

Stable Diffusion is a Natural Cross-Modal Decoder for Layered AI-generated Image Compression

2024-12-17 · Ruijie Chen, Qi Mao, Zhengxue Cheng

Recent advances in Artificial Intelligence Generated Content (AIGC) have garnered significant interest, accompanied by an increasing need to transmit and compress the vast number of AI-generated images (AIGIs). However, …

DecoderImage CompressionImage Generation

Interactive Semantic Featuring for Text Classification

2016-06-24 · Camille Jandot, Patrice Simard, Max Chickering, David Grangier 외

In text classification, dictionaries can be used to define human-comprehensible features. We propose an improvement to dictionary features called smoothed dictionary features. These features recognize document contexts i…

ClassificationGeneral Classificationtext-classificationText Classification

Semantic Compression via Multimodal Representation Learning

2025-09-29 · Eleonora Grassucci, Giordano Cicchetti, Aurelio Uncini, Danilo Comminiello arxiv

Multimodal representation learning produces high-dimensional embeddings that align diverse modalities in a shared latent space. While this enables strong generalization, it also introduces scalability challenges, both in…

Representation Learning

Enhancing the Comprehensibility of Text Explanations via Unsupervised Concept Discovery

2025-05-26 · Yifan Sun, Danding Wang, Qiang Sheng, Juan Cao 외

Concept-based explainable approaches have emerged as a promising method in explainable AI because they can interpret models in a way that aligns with human reasoning. However, their adaption in the text domain remains li…

Compression Beyond Pixels: Semantic Compression with Multimodal Foundation Models

2025-09-07 · Ruiqi Shen, Haotian Wu, Wenjing Zhang, Jiangjing Hu 외 arxiv

Recent deep learning-based methods for lossy image compression achieve competitive rate-distortion performance through extensive end-to-end training and advanced architectures. However, emerging applications increasingly…

Image Compression