VLF-MSC: Vision-Language Feature-Based Multimodal Semantic Communication System
We propose Vision-Language Feature-based Multimodal Semantic Communication (VLF-MSC), a unified system that transmits a single compact vision-language representation to support both image and text generation at the receiver. Unlike existing semantic communication techniques that process each modality separately, VLF-MSC employs a pre-trained vision-language model (VLM) to encode the source image into a vision-language semantic feature (VLF), which is transmitted over the wireless channel. At the receiver, a decoder-based language model and a diffusion-based image generator are both conditioned on the VLF to produce a descriptive text and a semantically aligned image. This unified representation eliminates the need for modality-specific streams or retransmissions, improving spectral efficiency and adaptability. By leveraging foundation models, the system achieves robustness to channel noise while preserving semantic fidelity. Experiments demonstrate that VLF-MSC outperforms text-only and image-only baselines, achieving higher semantic accuracy for both modalities under low SNR with significantly reduced bandwidth.
Code (0)
등록된 구현이 없습니다.
Tasks
Semantic CommunicationText GenerationSimilar Papers 제목 키워드 기반
Large Language Model-Driven Distributed Integrated Multimodal Sensing and Semantic Communications
Traditional single-modal sensing systems-based solely on either radio frequency (RF) or visual data-struggle to cope with the demands of complex and dynamic environments. Furthermore, single-device systems are constraine…
Language ModelingLanguage ModellingLarge Language ModelSemantic CommunicationTokenCom: Vision-Language Model for Multimodal and Multitask Token Communications
Visual-Language Models (VLMs), with their strong capabilities in image and text understanding, offer a solid foundation for intelligent communications. However, their effectiveness is constrained by limited token granula…
Image Generation with Supervised Selection Based on Multimodal Features for Semantic Communications
Semantic communication (SemCom) has emerged as a promising technique for the next-generation communication systems, in which the generation at the receiver side is allowed with semantic features' recovery. However, the m…
Image GenerationImage ReconstructionSemantic CommunicationTask-Oriented Semantic Communication in Large Multimodal Models-based Vehicle Networks
Task-oriented semantic communication has emerged as a fundamental approach for enhancing performance in various communication scenarios. While recent advances in Generative Artificial Intelligence (GenAI), such as Large …
Question AnsweringSemantic CommunicationVisual Question AnsweringVisual Question Answering (VQA)Task-Oriented Multi-User Semantic Communications for VQA Task
Semantic communications focus on the transmission of semantic features. In this letter, we consider a task-oriented multi-user semantic communication system for multimodal data transmission. Particularly, partial users t…
Question AnsweringSemantic CommunicationVisual Question AnsweringVisual Question Answering (VQA)