Feature Upsampling
1개 벤치마크 · 논문 39편 · 이 태스크의 논문 보기 →
Benchmarks
ImageNet
Most implemented
Deep Image Prior
CARAFE: Content-Aware ReAssembly of FEatures
LoftUp: Learning a Coordinate-Based Feature Upsampler for Vision Foundation Models
FeatUp: A Model-Agnostic Framework for Features at Any Resolution
On Point Affiliation in Feature Upsampling
SAPA: Similarity-Aware Point Affiliation for Feature Upsampling
Papers
RaysUp: Ultra-light Universal Feature Upsampling via Geometry-Aware Ray Representation
Pre-trained Vision Foundation Models (VFMs) have become central to modern computer vision due to their powerful semantic representations and strong generalization ability. However, their patchified or pooled outputs are …
Feature UpsamplingViT-Up: Faithful Feature Upsampling for Vision Transformers
Vision Transformers (ViTs) have become a dominant architecture for visual representation learning, providing exceptionally strong and broadly reusable backbone features. However, ViTs are commonly operated on relatively …
Semantic correspondenceRepresentation LearningSemantic SegmentationFeature UpsamplingWeighted Reverse Convolution for Feature Upsampling
Pre-trained vision foundation models (VFMs) provide strong semantic representations, yet their patch-level features are inherently coarse, limiting their effectiveness on tasks requiring fine-grained localization, dense …
Video Object SegmentationComputational EfficiencyFeature UpsamplingDepth EstimationDINO Soars: DINOv3 for Open-Vocabulary Semantic Segmentation of Remote Sensing Imagery
The remote sensing (RS) domain suffers from a lack of densely labeled datasets, which are costly to obtain. Thus, models that can segment RS imagery well without supervised fine-tuning are valuable, but existing solution…
Open Vocabulary Semantic SegmentationFeature UpsamplingFrozen Vision Transformers for Dense Prediction on Small Datasets: A Case Study in Arrow Localization
We present a system for automated detection, localization, and scoring of arrow punctures on 40\,cm indoor archery target faces, trained on only 48 annotated photographs (5{,}084 punctures). Our pipeline combines three c…
Feature UpsamplingHD-VGGT: High-Resolution Visual Geometry Transformer
High-resolution imagery is essential for accurate 3D reconstruction, as many geometric details only emerge at fine spatial scales. Recent feed-forward approaches, such as the Visual Geometry Grounded Transformer (VGGT), …
Feature Upsampling3D Reconstruction