paper-with-me

홈 › Papers

Disentangling 3D Prototypical Networks For Few-Shot Concept Learning

2020-11-06 · ICLR 2021 1 · Mihir Prabhudesai, Shamit Lal, Darshan Patil, Hsiao-Yu Tung, Adam W Harley, Katerina Fragkiadaki

We present neural architectures that disentangle RGB-D images into objects' shapes and styles and a map of the background scene, and explore their applications for few-shot 3D object detection and few-shot concept classification. Our networks incorporate architectural biases that reflect the image formation process, 3D geometry of the world scene, and shape-style interplay. They are trained end-to-end self-supervised by predicting views in static scenes, alongside a small number of 3D object boxes. Objects and scenes are represented in terms of 3D feature grids in the bottleneck of the network. We show that the proposed 3D neural representations are compositional: they can generate novel 3D scene feature maps by mixing object shapes and styles, resizing and adding the resulting object 3D feature maps over background scene feature maps. We show that classifiers for object categories, color, materials, and spatial relationships trained over the disentangled 3D feature sub-spaces generalize better with dramatically fewer examples than the current state-of-the-art, and enable a visual question answering system that uses them as its modules to generalize one-shot to novel objects in the scene.

📄 PDF Abstract BibTeX arXiv:2011.03367

Code (1)

mihirp1998/Disentangling-3D-Prototypical-Nets 공식 구현 pytorch

Tasks

3D geometry3D Object DetectionObjectobject-detectionObject DetectionQuestion AnsweringVisual Question AnsweringVisual Question Answering (VQA)

Similar Papers 제목 키워드 기반

Interpretable Few-Shot Image Classification via Prototypical Concept-Guided Mixture of LoRA Experts

2025-06-05 · Zhong Ji, Rongshuai Wei, Jingren Liu, Yanwei Pang 외

Self-Explainable Models (SEMs) rely on Prototypical Concept Learning (PCL) to enable their visual recognition processes more interpretable, but they often struggle in data-scarce settings where insufficient training samp…

Explainable ModelsFew-Shot Image Classificationimage-classificationImage Classification

Variational Prototyping-Encoder: One-Shot Learning with Prototypical Images

2019-04-17 · CVPR 2019 6 · Junsik Kim, Tae-Hyun Oh, Seokju Lee, Fei Pan 외

In daily life, graphic symbols, such as traffic signs and brand logos, are ubiquitously utilized around us due to its intuitive expression beyond language boundary. We tackle an open-set graphic symbol recognition proble…

Metric LearningOne-Shot LearningTranslation

Prototypical Priors: From Improving Classification to Zero-Shot Learning

2015-12-03 · Saumya Jetley, Bernardino Romera-Paredes, Sadeep Jayasumana, Philip Torr

Recent works on zero-shot learning make use of side information such as visual attributes or natural language semantics to define the relations between output visual classes and then use these relationships to draw infer…

ClassificationGeneral ClassificationZero-Shot Learning

MeDSLIP: Medical Dual-Stream Language-Image Pre-training for Fine-grained Alignment

2024-03-15 · Wenrui Fan, Mohammod Naimul Islam Suvon, Shuo Zhou, Xianyuan Liu 외

Vision-language pre-training (VLP) models have shown significant advancements in the medical domain. Yet, most VLP models align raw reports to images at a very coarse level, without modeling fine-grained relationships be…

AnatomyContrastive LearningRepresentation Learning

Prototypical Information Bottlenecking and Disentangling for Multimodal Cancer Survival Prediction

2024-01-03 · Yilan Zhang, Yingxue Xu, Jianqi Chen, Fengying Xie 외

Multimodal learning significantly benefits cancer survival prediction, especially the integration of pathological images and genomic data. Despite advantages of multimodal learning for cancer survival prediction, massive…

DisentanglementSurvival Predictionwhole slide images