paper-with-me

홈 › Papers

FeatSharp: Your Vision Model Features, Sharper

2025-02-22 · Mike Ranzinger, Greg Heinrich, Pavlo Molchanov, Jan Kautz, Bryan Catanzaro, Andrew Tao

The feature maps of vision encoders are fundamental to myriad modern AI tasks, ranging from core perception algorithms (e.g. semantic segmentation, object detection, depth perception, etc.) to modern multimodal understanding in vision-language models (VLMs). Currently, in computer vision, the frontier of general purpose vision backbones are Vision Transformers (ViT), typically trained using contrastive loss (e.g. CLIP). A key problem with most off-the-shelf ViTs, particularly CLIP, is that these models are inflexibly low resolution. Most run at 224x224px, while the "high resolution" versions are around 378-448px, but still inflexible. We introduce a novel method to coherently and cheaply upsample the feature maps of low-res vision encoders while picking up on fine-grained details that would otherwise be lost due to resolution. We demonstrate the effectiveness of this approach on core perception tasks as well as within agglomerative model (RADIO) training as a way of providing richer targets for distillation.

📄 PDF Abstract BibTeX arXiv:2502.16025

Code (1)

nvlabs/radio pytorch

Tasks

modelobject-detectionObject DetectionSemantic Segmentation

Methods 이 논문이 사용한 방법론

CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Bootstrap Your Own Prior: Towards Distribution-Agnostic Novel Class Discovery

2023-01-01 · CVPR 2023 1 · Muli Yang, Liancheng Wang, Cheng Deng, Hanwang Zhang

Novel Class Discovery (NCD) aims to discover unknown classes without any annotation, by exploiting the transferable knowledge already learned from a base set of known classes. Existing works hold an impractical assum…

Novel Class Discovery

High- and Low-level image component decomposition using VAEs for improved reconstruction and anomaly detection

2019-11-27 · David Zimmerer, Jens Petersen, Klaus Maier-Hein

Variational Auto-Encoders have often been used for unsupervised pretraining, feature extraction and out-of-distribution and anomaly detection in the medical field. However, VAEs often lack the ability to produce sharp im…

Anomaly DetectionOut of Distribution (OOD) Detection

Muizalix Pro

2025-02-07 · 02/07 2025 2 · Muizalix Pro

At Muizalix, we’re a dedicated team of creative professionals and digital experts passionate about helping businesses grow. Specializing in social media marketing, SEO, content creation, and website design, we provide co…

Marketing

Hand Gesture Real Time Paint Tool - Box

2017-09-03 · Vandit Gajjar, Viraj Mavani, Ayesha Gurnani

With current development universally in computing, now a days user interaction approaches with mouse, keyboard, touch-pens etc. are not sufficient. Directly using of hands or hand gestures as an input device is a method …

BIG-bench Machine Learning

ChromaCorrect: Prescription Correction in Virtual Reality Headsets through Perceptual Guidance

2022-12-08 · Ahmet Güzel, Jeanne Beyazian, PRANEETH CHAKRAVARTHULA, Kaan Akşit

A large portion of today's world population suffer from vision impairments and wear prescription eyeglasses. However, eyeglasses causes additional bulk and discomfort when used with augmented and virtual reality headsets…