paper-with-me

Papers

PAN-Crafter: Learning Modality-Consistent Alignment for PAN-Sharpening

2025-05-29 · Jeonghyeok Do, Sungpyo Kim, Geunhyuk Youk, Jaehyup Lee, Munchurl Kim

PAN-sharpening aims to fuse high-resolution panchromatic (PAN) images with low-resolution multi-spectral (MS) images to generate high-resolution multi-spectral (HRMS) outputs. However, cross-modality misalignment -- caused by sensor placement, acquisition timing, and resolution disparity -- induces a fundamental challenge. Conventional deep learning methods assume perfect pixel-wise alignment and rely on per-pixel reconstruction losses, leading to spectral distortion, double edges, and blurring when misalignment is present. To address this, we propose PAN-Crafter, a modality-consistent alignment framework that explicitly mitigates the misalignment gap between PAN and MS modalities. At its core, Modality-Adaptive Reconstruction (MARs) enables a single network to jointly reconstruct HRMS and PAN images, leveraging PAN's high-frequency details as auxiliary self-supervision. Additionally, we introduce Cross-Modality Alignment-Aware Attention (CM3A), a novel mechanism that bidirectionally aligns MS texture to PAN structure and vice versa, enabling adaptive feature refinement across modalities. Extensive experiments on multiple benchmark datasets demonstrate that our PAN-Crafter outperforms the most recent state-of-the-art method in all metrics, even with 50.11$\times$ faster inference time and 0.63$\times$ the memory size. Furthermore, it demonstrates strong generalization performance on unseen satellite datasets, showing its robustness across different conditions.

📄 PDF Abstract BibTeX arXiv:2505.23367

Code (0)

등록된 구현이 없습니다.

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음

Similar Papers 제목 키워드 기반

FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

2024-07-01 · Yiming Zhang, Yicheng Gu, Yanhong Zeng, Zhening Xing 외

We study Neural Foley, the automatic generation of high-quality sound effects synchronizing with videos, enabling an immersive audio-visual experience. Despite its wide range of applications, existing approaches encounte…

Audio GenerationVideo AlignmentVideo Synchronization

SceneCrafter: Controllable Multi-View Driving Scene Editing

2025-01-01 · CVPR 2025 1 · Zehao Zhu, Yuliang Zou, Chiyu Max Jiang, Bo Sun 외

Simulation is crucial for developing and evaluating autonomous vehicle (AV) systems. Recent literature builds on a new generation of generative models to synthesize highly realistic images for full-stack simulation. …

Diffusion-Sharpening: Fine-tuning Diffusion Models with Denoising Trajectory Sharpening

2025-02-17 · Ye Tian, Ling Yang, Xinchen Zhang, Yunhai Tong 외

We propose Diffusion-Sharpening, a fine-tuning approach that enhances downstream alignment by optimizing sampling trajectories. Existing RL-based fine-tuning methods focus on single training timesteps and neglect traject…

Denoising

ColonCrafter: A Depth Estimation Model for Colonoscopy Videos Using Diffusion Priors

2025-09-16 · Romain Hardy, Tyler Berzin, Pranav Rajpurkar arxiv

Three-dimensional (3D) scene understanding in colonoscopy presents significant challenges that necessitate automated methods for accurate depth estimation. However, existing depth estimation models for endoscopy struggle…

Point Cloud GenerationScene Understanding3D ReconstructionDepth Estimation

TextCrafter: Accurately Rendering Multiple Texts in Complex Visual Scenes

2025-03-30 · Nikai Du, Zhennan Chen, Zhizhou Chen, Shan Gao 외

This paper explores the task of Complex Visual Text Generation (CVTG), which centers on generating intricate textual content distributed across diverse regions within visual images. In CVTG, image generation models often…

2kImage GenerationText Generation