paper-with-me

홈 › Papers

Semantic Jitter: Dense Supervision for Visual Comparisons via Synthetic Images

2016-12-19 · ICCV 2017 10 · Aron Yu, Kristen Grauman

Distinguishing subtle differences in attributes is valuable, yet learning to make visual comparisons remains non-trivial. Not only is the number of possible comparisons quadratic in the number of training images, but also access to images adequately spanning the space of fine-grained visual differences is limited. We propose to overcome the sparsity of supervision problem via synthetically generated images. Building on a state-of-the-art image generation engine, we sample pairs of training images exhibiting slight modifications of individual attributes. Augmenting real training image pairs with these examples, we then train attribute ranking models to predict the relative strength of an attribute in novel pairs of real images. Our results on datasets of faces and fashion images show the great promise of bootstrapping imperfect image generators to counteract sample sparsity for learning to rank.

📄 PDF Abstract BibTeX arXiv:1612.06341

Code (0)

등록된 구현이 없습니다.

Tasks

AttributeImage GenerationLearning-To-Rank

Similar Papers 제목 키워드 기반

From Dense Prediction to Visual Editing: Structured Supervision for Unified Image and Video Creation

2026-08-13 · Zhefan Rao, Bin Zou, Xuanhua He, Chong Hou Choi 외 arxiv

Unified image and video creation requires a model to follow diverse instructions while preserving identity, geometry, and temporal structure from visual context. However, semantic-only conditioning and creation-only trai…

Enhancing 3D LiDAR Segmentation by Shaping Dense and Accurate 2D Semantic Predictions

2026-02-21 · Xiaoyu Dong, Tiankui Xian, Wanshui Gan, Naoto Yokoya arxiv

Semantic segmentation of 3D LiDAR point clouds is important in urban remote sensing for understanding real-world street environments. This task, by projecting LiDAR point clouds and 3D semantic labels as sparse maps, can…

Semantic SegmentationPoint Clouds

Separating the "Chirp" from the "Chat": Self-supervised Visual Grounding of Sound and Language

2024-06-09 · CVPR 2024 1 · Mark Hamilton, Andrew Zisserman, John R. Hershey, William T. Freeman

We present DenseAV, a novel dual encoder grounding architecture that learns high-resolution, semantically meaningful, and audio-visually aligned features solely through watching videos. We show that DenseAV can discover …

Contrastive LearningCross-Modal RetrievalSemantic SegmentationSound Prompted Semantic Segmentation+2

Learning Multi-Modal Prototypes for Cross-Domain Few-Shot Object Detection

2026-02-21 · Wanqi Wang, Jingcai Guo, Yuxiang Cai, Zhi Chen arxiv

Cross-Domain Few-Shot Object Detection (CD-FSOD) aims to detect novel classes in unseen target domains given only a few labeled examples. While open-vocabulary detectors built on vision-language models (VLMs) transfer we…

Cross-Domain Few-Shot Object Detection

Revisiting Multi-Task Visual Representation Learning

2026-01-20 · Shangzhe Di, Zhonghua Zhai, Weidi Xie arxiv

Current visual representation learning remains bifurcated: vision-language models (e.g., CLIP) excel at global semantic alignment but lack spatial precision, while self-supervised methods (e.g., MAE, DINO) capture intric…

Representation LearningMulti-Task LearningSpatial Reasoning