paper-with-me

홈 › Papers

Temporal-Spatial Tubelet Embedding for Cloud-Robust MSI Reconstruction using MSI-SAR Fusion: A Multi-Head Self-Attention Video Vision Transformer Approach

2025-12-10 · Yiqun Wang, Lujun Li, Meiru Yue, Radu State arxiv

Cloud cover in multispectral imagery (MSI) significantly hinders early-season crop mapping by corrupting spectral information. Existing Vision Transformer(ViT)-based time-series reconstruction methods, like SMTS-ViT, often employ coarse temporal embeddings that aggregate entire sequences, causing substantial information loss and reducing reconstruction accuracy. To address these limitations, a Video Vision Transformer (ViViT)-based framework with temporal-spatial fusion embedding for MSI reconstruction in cloud-covered regions is proposed in this study. Non-overlapping tubelets are extracted via 3D convolution with constrained temporal span $(t=2)$, ensuring local temporal coherence while reducing cross-day information degradation. Both MSI-only and SAR-MSI fusion scenarios are considered during the experiments. Comprehensive experiments on 2020 Traill County data demonstrate notable performance improvements: MTS-ViViT achieves a 2.23\% reduction in MSE compared to the MTS-ViT baseline, while SMTS-ViViT achieves a 10.33\% improvement with SAR integration over the SMTS-ViT baseline. The proposed framework effectively enhances spectral reconstruction quality for robust agricultural monitoring.

📄 PDF Abstract BibTeX arXiv:2512.09471

Code (0)

등록된 구현이 없습니다.

Tasks

Spectral Reconstruction

Similar Papers 제목 키워드 기반

In Defense of Clip-based Video Relation Detection

2023-07-18 · Meng Wei, Long Chen, Wei Ji, Xiaoyu Yue 외

Video Visual Relation Detection (VidVRD) aims to detect visual relationship triplets in videos using spatial bounding boxes and temporal boundaries. Existing VidVRD methods can be broadly categorized into bottom-up and t…

Feature CompressionObject TrackingRelationVideo Visual Relation Detection

Spatiotemporal Learning with Context-aware Video Tubelets for Ultrasound Video Analysis

2025-03-21 · Gary Y. Li, Li Chen, Bryson Hicks, Nikolai Schnittke 외

Computer-aided pathology detection algorithms for video-based imaging modalities must accurately interpret complex spatiotemporal information by integrating findings across multiple frames. Current state-of-the-art metho…

object-detectionObject DetectionVideo Classification

Video-based Human-Object Interaction Detection from Tubelet Tokens

2022-06-04 · Danyang Tu, Wei Sun, Xiongkuo Min, Guangtao Zhai 외

We present a novel vision Transformer, named TUTOR, which is able to learn tubelet tokens, served as highly-abstracted spatiotemporal representations, for video-based human-object interaction (V-HOI) detection. The tubel…

Human-Object Interaction Detection

Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution

2024-03-25 · CVPR 2024 1 · Zhikai Chen, Fuchen Long, Zhaofan Qiu, Ting Yao 외

Diffusion models are just at a tipping point for image super-resolution task. Nevertheless, it is not trivial to capitalize on diffusion models for video super-resolution which necessitates not only the preservation of v…

DecoderDenoisingImage Super-ResolutionSuper-Resolution+3

Object Detection in Videos by Short and Long Range Object Linking

2018-01-30 · IEEE Transactions on Pattern Analysis and Machine Intelligence(TPAM) 2018 1 · Peng Tang † Chunyu Wang ‡ Xinggang Wang † Wenyu Liu † Wenjun Zeng ‡ Jingdong Wang ‡ † School of EIC, Huazhong University of Science and Technology   ‡ Microsoft Research Asia

We address the problem of detecting objects in videos with the interest in exploring temporal contexts. Our core idea is to link objects in the short and long ranges for improving the classification quality. Our approach…

ClassificationObjectobject-detectionObject Detection