Multi-Modal Semantic Communication
Semantic communication aims to transmit information most relevant to a task rather than raw data, offering significant gains in communication efficiency for applications such as telepresence, augmented reality, and remote sensing. Recent transformer-based approaches have used self-attention maps to identify informative regions within images, but they often struggle in complex scenes with multiple objects, where self-attention lacks explicit task guidance. To address this, we propose a novel Multi-Modal Semantic Communication framework that integrates text-based user queries to guide the information extraction process. Our proposed system employs a cross-modal attention mechanism that fuses visual features with language embeddings to produce soft relevance scores over the visual data. Based on these scores and the instantaneous channel bandwidth, we use an algorithm to transmit image patches at adaptive resolutions using independently trained encoder-decoder pairs, with total bitrate matching the channel capacity. At the receiver, the patches are reconstructed and combined to preserve task-critical information. This flexible and goal-driven design enables efficient semantic communication in complex and bandwidth-constrained environments.
Code (0)
등록된 구현이 없습니다.
Tasks
Information ExtractionSemantic CommunicationSimilar Papers 제목 키워드 기반
Multi-Modal Fusion-Based Multi-Task Semantic Communication System
In recent years, there has been significant progress in semantic communication systems empowered by deep learning techniques. It has greatly improved the efficiency of information transmission. Nevertheless, traditional …
Semantic CommunicationRate-Adaptive Coding Mechanism for Semantic Communications With Multi-Modal Data
Recently, the ever-increasing demand for bandwidth in multi-modal communication systems requires a paradigm shift. Powered by deep learning, semantic communications are applied to multi-modal scenarios to boost communica…
DecoderSemantic CommunicationMulti-Modal Self-Supervised Semantic Communication
Semantic communication is emerging as a promising paradigm that focuses on the extraction and transmission of semantic meanings using deep learning techniques. While current research primarily addresses the reduction of …
Self-Supervised LearningSemantic CommunicationPilot-guided Multimodal Semantic Communication for Audio-Visual Event Localization
Multimodal semantic communication, which integrates various data modalities such as text, images, and audio, significantly enhances communication efficiency and reliability. It has broad application prospects in fields s…
audio-visual event localizationAutonomous DrivingSemantic CommunicationTask-Oriented Multi-User Semantic Communications for VQA Task
Semantic communications focus on the transmission of semantic features. In this letter, we consider a task-oriented multi-user semantic communication system for multimodal data transmission. Particularly, partial users t…
Question AnsweringSemantic CommunicationVisual Question AnsweringVisual Question Answering (VQA)