paper-with-me

홈 › Papers

Dense Cross-Scale Image Alignment With Fully Spatial Correlation and Just Noticeable Difference Guidance

2025-11-12 · Jinkun You, Jiaxue Li, Jie Zhang, Yicong Zhou arxiv

Existing unsupervised image alignment methods exhibit limited accuracy and high computational complexity. To address these challenges, we propose a dense cross-scale image alignment model. It takes into account the correlations between cross-scale features to decrease the alignment difficulty. Our model supports flexible trade-offs between accuracy and efficiency by adjusting the number of scales utilized. Additionally, we introduce a fully spatial correlation module to further improve accuracy while maintaining low computational costs. We incorporate the just noticeable difference to encourage our model to focus on image regions more sensitive to distortions, eliminating noticeable alignment errors. Extensive quantitative and qualitative experiments demonstrate that our method surpasses state-of-the-art approaches.

📄 PDF Abstract BibTeX arXiv:2511.09028

Code (0)

등록된 구현이 없습니다.

Similar Papers 제목 키워드 기반

MUSE: Multi-Scale Dense Self-Distillation for Nucleus Detection and Classification

2025-11-07 · Zijiang Yang, Hanqing Chao, Bokai Zhao, Yelin Yang 외 arxiv

Nucleus detection and classification (NDC) in histopathology analysis is a fundamental task that underpins a wide range of high-level pathology applications. However, existing methods heavily rely on labor-intensive nucl…

Self-Supervised Learning

Incorporating Dense Knowledge Alignment into Unified Multimodal Representation Models

2025-01-01 · CVPR 2025 1 · Yuhao Cui, Xinxing Zu, Wenhua Zhang, Zhongzhou Zhao 외

Leveraging Large Language Models (LLMs) for text representation has achieved significant success, but the exploration of using Multimodal LLMs (MLLMs) for multimodal representation remains limited. Previous MLLM-base…

Contrastive LearningCross-Modal RetrievalRetrieval

Tango3D: Towards Alignment for Global and Local 2D-3D Correspondence

2026-05-19 · Zebin He, Mingxin Yang, Shuhui Yang, Hanxiao Sun 외 arxiv

Existing 3D foundation models typically align point clouds to frozen vision-language spaces like CLIP, which achieve strong cross-modal retrieval by compressing 3D shape into a global vector. However, this global-only al…

Cross-Modal RetrievalPoint Clouds

SegFly: A Dataset and 2D-3D-2D Paradigm for Aerial RGB-Thermal Semantic Segmentation at Scale

2026-03-18 · Markus Gross, Sai Bharadhwaj Matha, Rui Song, Viswanathan Muthuveerappan 외 arxiv

Semantic segmentation for uncrewed aerial vehicles (UAVs) is fundamental for aerial scene understanding, yet existing RGB and RGB-T datasets remain limited in scale, diversity, and annotation efficiency due to the high c…

Semantic SegmentationScene UnderstandingImage Registration

MDU-Net: Multi-scale Densely Connected U-Net for biomedical image segmentation

2018-12-02 · Jiawei Zhang, Yuzhen Jin, Jilan Xu, Xiaowei Xu 외

Radiologist is "doctor's doctor", biomedical image segmentation plays a central role in quantitative analysis, clinical diagnosis, and medical intervention. In the light of the fully convolutional networks (FCN) and U-Ne…

DecoderImage SegmentationQuantizationSegmentation+1