paper-with-me

Papers

UniStitch: Unifying Semantic and Geometric Features for Image Stitching

2026-03-11 · Yuan Mei, Lang Nie, Kang Liao, Yunqiu Xu, Chunyu Lin, Bin Xiao arxiv

Traditional image stitching methods estimate warps from hand-crafted geometric features, whereas recent learning-based solutions leverage semantic features from neural networks instead. These two lines of research have largely diverged along separate evolution, with virtually no meaningful convergence to date. In this paper, we take a pioneering step to bridge this gap by unifying semantic and geometric features with UniStitch, a unified image stitching framework from multimodal features. To align discrete geometric features (i.e., keypoint) with continuous semantic feature maps, we present a Neural Point Transformer (NPT) module, which transforms unordered, sparse 1D geometric keypoints into ordered, dense 2D semantic maps. Then, to integrate the advantages of both representations, an Adaptive Mixture of Experts (AMoE) module is designed to fuse geometric and semantic representations. It dynamically shifts focus toward more reliable features during the fusion process, allowing the model to handle complex scenes, especially when either modality might be compromised. The fused representation can be adopted into common deep stitching pipelines, delivering significant performance gains over any single feature. Experiments show that UniStitch outperforms existing state-of-the-art methods with a large margin, paving the way for a unified paradigm between traditional and learning-based image stitching.

📄 PDF Abstract BibTeX arXiv:2603.10568

Code (0)

등록된 구현이 없습니다.

Tasks

Image Stitching

Similar Papers 제목 키워드 기반

Cross-Dataset Semantic Segmentation Performance Analysis: Unifying NIST Point Cloud City Datasets for 3D Deep Learning

2025-08-01 · Alexander Nikitas Dimopoulos, Joseph Grasso arxiv

This study analyzes semantic segmentation performance across heterogeneously labeled point-cloud datasets relevant to public safety applications, including pre-incident planning systems derived from lidar scans. Using NI…

Semantic Segmentation

Thinking with Spatial Code for Physical-World Video Reasoning

2026-03-05 · Jieneng Chen, Wenxin Ma, Ruisheng Yuan, Yunzhi Zhang 외 arxiv

We introduce Thinking with Spatial Code, a framework that transforms RGB video into explicit, temporally coherent 3D representations for physical-world visual question answering. We highlight the empirical finding that o…

Visual Question AnsweringReinforcement Learning

UniOcc: Unifying Vision-Centric 3D Occupancy Prediction with Geometric and Semantic Rendering

2023-06-15 · Mingjie Pan, Li Liu, Jiaming Liu, Peixiang Huang 외

In this technical report, we present our solution, named UniOCC, for the Vision-Centric 3D occupancy prediction track in the nuScenes Open Dataset Challenge at CVPR 2023. Existing methods for occupancy prediction primari…

PredictionPrediction Of Occupancy Grid Maps

GASP: Unifying Geometric and Semantic Self-Supervised Pre-training for Autonomous Driving

2025-03-19 · William Ljungbergh, Adam Lilja, Adam Tonderski. Arvid Laveno Ling, Carl Lindström 외

Self-supervised pre-training based on next-token prediction has enabled large language models to capture the underlying structure of text, and has led to unprecedented performance on a large array of tasks when applied a…

Autonomous DrivingTrajectory Prediction

Unifying Geometric Features and Facial Action Units for Improved Performance of Facial Expression Analysis

2016-06-02 · Mehdi Ghayoumi, Arvind K Bansal

Previous approaches to model and analyze facial expression analysis use three different techniques: facial action units, geometric features and graph based modelling. However, previous approaches have treated these techn…

Dimensionality ReductionEmotion RecognitionGeneral Classificationimage-classification+1