paper-with-me

Papers

SemDINO: Foundation Prior-Guided Cross-Temporal Semantic Alignment Network for Remote Sensing Change Detection

2026-06-08 · Xinyu Tong, Meihua Zhou, Jinxiao Sun, Yingjie Tang, Lei Wang arxiv

Semantic change detection (SCD) in remote sensing aims to identify land-cover transitions between bi-temporal observations while suppressing pseudo-changes caused by illumination variations, seasonal differences, and registration errors. Although Vision Foundation Models (VFMs) provide transferable semantic priors, their application to SCD remains challenging due to the mismatch between foundation-model representations and task-specific spatial features, as well as temporal-order sensitivity. To address these issues, this paper proposes SemDINO, a foundation prior-guided framework that integrates transferable vision foundation model priors with hierarchical convolutional representations for cross-temporal semantic reasoning. Specifically, a Gated Pyramid Fusion (PyFu) module is developed to adaptively combine foundation-model semantics with CNN spatial details while reducing domain noise. A Multi-scale Temporal Bi-directional Transformer (M-TBTT) is introduced to achieve symmetric cross-temporal feature interaction and alleviate temporal-order bias. Furthermore, a Feature Change Enhancement (FeaCE) flow is designed to refine aligned representations and distinguish genuine semantic transitions from pseudo variations. Finally, a multi-branch decoupled prediction head jointly generates change masks, bi-temporal semantic maps, and edge constraints. Extensive experiments across five benchmark datasets demonstrate that SemDINO consistently outperforms state-of-the-art methods on both semantic and binary change detection tasks. The results validate the effectiveness of alignment-oriented representation learning for robust remote sensing change analysis.

📄 PDF Abstract BibTeX arXiv:2606.09772

Code (0)

등록된 구현이 없습니다.

Tasks

Change Detection

Similar Papers 제목 키워드 기반

Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animation

2025-12-05 · Su Sun, Cheng Zhao, Himangi Mittal, Gaurav Mittal 외 arxiv

Generating dynamic 4D objects from sparse inputs is difficult because it demands joint preservation of appearance and motion coherence across views and time while suppressing artifacts and temporal drift. We hypothesize …

Video Generation

Temporal-Guided Visual Foundation Models for Event-Based Vision

2025-11-09 · Ruihao Xia, Junhong Cai, Luziwei Leng, Liuyi Wang 외 arxiv

Event cameras offer unique advantages for vision tasks in challenging environments, yet processing asynchronous event streams remains an open challenge. While existing methods rely on specialized architectures or resourc…

Semantic SegmentationEvent-based visionDepth EstimationObject Detection

UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors

2026-03-16 · Xiaoyan Cong, Zekun Li, Zhiyang Dou, Hongyu Li 외 arxiv

Large-scale foundation models (LFMs) have recently made impressive progress in text-to-motion generation by learning strong generative priors from massive 3D human motion datasets and paired text descriptions. However, h…

Motion Synthesis

Depth Completion as Parameter-Efficient Test-Time Adaptation

2026-02-16 · Bingxin Ke, Qunjie Zhou, Jiahui Huang, Xuanchi Ren 외 arxiv

We introduce CAPA, a parameter-efficient test-time optimization framework that adapts pre-trained 3D foundation models (FMs) for depth completion, using sparse geometric cues. Unlike prior methods that train task-specifi…

parameter-efficient fine-tuningTest-time AdaptationDepth Completion

BrainRVQ: A High-Fidelity EEG Foundation Model via Dual-Domain Residual Quantization and Hierarchical Autoregression

2026-02-18 · Mingzhe Cui, Tao Chen, Yang Jiao, Yiqin Wang 외 arxiv

Developing foundation models for electroencephalography (EEG) remains challenging due to the signal's low signal-to-noise ratio and complex spectro-temporal non-stationarity. Existing approaches often overlook the hierar…