paper-with-me

홈 › Papers

VLF-MSC: Vision-Language Feature-Based Multimodal Semantic Communication System

2025-11-13 · Gwangyeon Ahn, Jiwan Seo, Joonhyuk Kang arxiv

We propose Vision-Language Feature-based Multimodal Semantic Communication (VLF-MSC), a unified system that transmits a single compact vision-language representation to support both image and text generation at the receiver. Unlike existing semantic communication techniques that process each modality separately, VLF-MSC employs a pre-trained vision-language model (VLM) to encode the source image into a vision-language semantic feature (VLF), which is transmitted over the wireless channel. At the receiver, a decoder-based language model and a diffusion-based image generator are both conditioned on the VLF to produce a descriptive text and a semantically aligned image. This unified representation eliminates the need for modality-specific streams or retransmissions, improving spectral efficiency and adaptability. By leveraging foundation models, the system achieves robustness to channel noise while preserving semantic fidelity. Experiments demonstrate that VLF-MSC outperforms text-only and image-only baselines, achieving higher semantic accuracy for both modalities under low SNR with significantly reduced bandwidth.

📄 PDF Abstract BibTeX arXiv:2511.10074

Code (0)

등록된 구현이 없습니다.

Tasks

Semantic CommunicationText Generation

Similar Papers 제목 키워드 기반

Large Language Model-Driven Distributed Integrated Multimodal Sensing and Semantic Communications

2025-05-20 · Yubo Peng, Luping Xiang, Bingxin Zhang, Kun Yang

Traditional single-modal sensing systems-based solely on either radio frequency (RF) or visual data-struggle to cope with the demands of complex and dynamic environments. Furthermore, single-device systems are constraine…

Language ModelingLanguage ModellingLarge Language ModelSemantic Communication

TokenCom: Vision-Language Model for Multimodal and Multitask Token Communications

2026-02-28 · Feibo Jiang, Siwei Tu, Li Dong, Xiaolong Li 외 arxiv

Visual-Language Models (VLMs), with their strong capabilities in image and text understanding, offer a solid foundation for intelligent communications. However, their effectiveness is constrained by limited token granula…

Image Generation with Supervised Selection Based on Multimodal Features for Semantic Communications

2024-11-26 · Chengyang Liang, Dong Li

Semantic communication (SemCom) has emerged as a promising technique for the next-generation communication systems, in which the generation at the receiver side is allowed with semantic features' recovery. However, the m…

Image GenerationImage ReconstructionSemantic Communication

Task-Oriented Semantic Communication in Large Multimodal Models-based Vehicle Networks

2025-05-05 · Baoxia Du, Hongyang Du, Dusit Niyato, Ruidong Li

Task-oriented semantic communication has emerged as a fundamental approach for enhancing performance in various communication scenarios. While recent advances in Generative Artificial Intelligence (GenAI), such as Large …

Question AnsweringSemantic CommunicationVisual Question AnsweringVisual Question Answering (VQA)

Task-Oriented Multi-User Semantic Communications for VQA Task

2021-08-16 · Huiqiang Xie, Zhijin Qin, Geoffrey Ye Li

Semantic communications focus on the transmission of semantic features. In this letter, we consider a task-oriented multi-user semantic communication system for multimodal data transmission. Particularly, partial users t…

Question AnsweringSemantic CommunicationVisual Question AnsweringVisual Question Answering (VQA)