A Multi-Scale Spatial-Temporal Network for Wireless Video Transmission
Deep joint source-channel coding (DeepJSCC) has shown promise in wireless transmission of text, speech, and images within the realm of semantic communication. However, wireless video transmission presents greater challenges due to the difficulty of extracting and compactly representing both spatial and temporal features, as well as its significant bandwidth and computational resource requirements. In response, we propose a novel video DeepJSCC (VDJSCC) approach to enable end-to-end video transmission over a wireless channel. Our approach involves the design of a multi-scale vision Transformer encoder and decoder to effectively capture spatial-temporal representations over long-term frames. Additionally, we propose a dynamic token selection module to mask less semantically important tokens from spatial or temporal dimensions, allowing for content-adaptive variable-length video coding by adjusting the token keep ratio. Experimental results demonstrate the effectiveness of our VDJSCC approach compared to digital schemes that use separate source and channel codes, as well as other DeepJSCC schemes, in terms of reconstruction quality and bandwidth reduction.
Code (0)
등록된 구현이 없습니다.
Tasks
DecoderSemantic CommunicationMethods 이 논문이 사용한 방법론
Similar Papers 제목 키워드 기반
Social Relation Recognition From Videos via Multi-Scale Spatial-Temporal Reasoning
Discovering social relations, e.g., kinship, friendship, etc., from visual contents can make machines better interpret the behaviors and emotions of human beings. Existing studies mainly focus on recognizing social relat…
RelationTemporal-Spatial Feature Pyramid for Video Saliency Detection
Multi-level features are important for saliency detection. Better combination and use of multi-level features with time information can greatly improve the accuracy of the video saliency model. In order to fully combine …
DecoderSaliency DetectionVideo Saliency DetectionMotion Compensated Frequency Selective Extrapolation for Error Concealment in Video Coding
Although wireless and IP-based access to video content gives a new degree of freedom to the viewers, the risk of severe block losses caused by transmission errors is always present. The purpose of this paper is to presen…
Spatial-Temporal Correlation and Topology Learning for Person Re-Identification in Videos
Video-based person re-identification aims to match pedestrians from video sequences across non-overlapping camera views. The key factor for video person re-identification is to effectively exploit both spatial and tempor…
Person Re-IdentificationVideo-Based Person Re-IdentificationVideo DeinterlacingRegion-Based Multiscale Spatiotemporal Saliency for Video
Detecting salient objects from a video requires exploiting both spatial and temporal knowledge included in the video. We propose a novel region-based multiscale spatiotemporal saliency detection method for videos, where …
Saliency Detection