paper-with-me

Papers

Image Generation with Supervised Selection Based on Multimodal Features for Semantic Communications

2024-11-26 · Chengyang Liang, Dong Li

Semantic communication (SemCom) has emerged as a promising technique for the next-generation communication systems, in which the generation at the receiver side is allowed with semantic features' recovery. However, the majority of existing research predominantly utilizes a singular type of semantic information, such as text, images, or speech, to supervise and choose the generated source signals, which may not sufficiently encapsulate the comprehensive and accurate semantic information, and thus creating a performance bottleneck. In order to bridge this gap, in this paper, we propose and investigate a SemCom framework using multimodal information to supervise the generated image. To be specific, in this framework, we first extract semantic features at both the image and text levels utilizing the Convolutional Neural Network (CNN) architecture and the Contrastive Language-Image Pre-Training (CLIP) model before transmission. Then, we employ a generative diffusion model at the receiver to generate multiple images. In order to ensure the accurate extraction and facilitate high-fidelity image reconstruction, we select the "best" image with the minimum reconstruction errors by taking both the aided image and text semantic features into account. We further extend multimodal semantic communication (MMSemCom) system to the multiuser scenario for orthogonal transmission. Experimental results demonstrate that the proposed framework can not only achieve the enhanced fidelity and robustness in image transmission compared with existing communication systems but also sustain a high performance in the low signal-to-noise ratio (SNR) conditions.

📄 PDF Abstract BibTeX arXiv:2411.17428

Code (0)

등록된 구현이 없습니다.

Tasks

Image GenerationImage ReconstructionSemantic Communication

Methods 이 논문이 사용한 방법론

Diffusion Diffusion models generate samples by gradually removing noise from a signal, and their training objective can be expressed as a reweighted variational lower-bound…

Similar Papers 제목 키워드 기반

Seek Common Ground While Reserving Differences: Semi-Supervised Image-Text Sentiment Recognition

2025-01-01 · CVPR 2025 1 · Wuyou Xia, Guoli Jia, Sicheng Zhao, Jufeng Yang

Multimodal sentiment analysis has attracted extensive research attention as increasing users share images and texts to express their emotions and opinions on social media. Collecting large amounts of labeled sentimen…

Multimodal Sentiment AnalysisPseudo LabelPseudo Label FilteringSentiment Analysis

AutoCut: End-to-end advertisement video editing based on multimodal discretization and controllable generation

2026-03-30 · Milton Zhou, Sizhong Qin, Yongzhi Li, Quan Chen 외 arxiv

Short-form videos have become a primary medium for digital advertising, requiring scalable and efficient content creation. However, current workflows and AI tools remain disjoint and modality-specific, leading to high pr…

Story-oriented Image Selection and Placement

2019-09-02 · Sreyasi Nag Chowdhury, Simon Razniewski, Gerhard Weikum

Multimodal contents have become commonplace on the Internet today, manifested as news articles, social media posts, and personal or business blog posts. Among the various kinds of media (images, videos, graphics, icons, …

ArticlesCombinatorial OptimizationObject Recognition

EVLP:Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning

2025-11-03 · Xinyan Cai, Shiguang Wu, Dafeng Chi, Yuzheng Zhuang 외 arxiv

In complex embodied long-horizon manipulation tasks, effective task decomposition and execution require synergistic integration of textual logical reasoning and visual-spatial imagination to ensure efficient and accurate…

multimodal generationLogical Reasoning

ViKL: A Mammography Interpretation Framework via Multimodal Aggregation of Visual-knowledge-linguistic Features

2024-09-24 · Xin Wei, Yaling Tao, Changde Du, Gangming Zhao 외

Mammography is the primary imaging tool for breast cancer diagnosis. Despite significant strides in applying deep learning to interpret mammography images, efforts that focus predominantly on visual features often strugg…

Contrastive Learning