paper-with-me

홈 › Papers

Fine-Grained Visual Comparisons with Local Learning

2014-06-01 · CVPR 2014 6 · Aron Yu, Kristen Grauman

Given two images, we want to predict which exhibits a particular visual attribute more than the other---even when the two images are quite similar. Existing relative attribute methods rely on global ranking functions; yet rarely will the visual cues relevant to a comparison be constant for all data, nor will humans' perception of the attribute necessarily permit a global ordering. To address these issues, we propose a local learning approach for fine-grained visual comparisons. Given a novel pair of images, we learn a local ranking model on the fly, using only analogous training comparisons. We show how to identify these analogous pairs using learned metrics. With results on three challenging datasets -- including a large newly curated dataset for fine-grained comparisons -- our method outperforms state-of-the-art methods for relative attribute prediction.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Attribute

Similar Papers 제목 키워드 기반

FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting

2026-04-30 · Fengxian Ji, Jingpu Yang, Zirui Song, Yuanxi Wang 외 arxiv

Despite the rapid progress of large vision-language models (LVLMs), fine-grained, state-conditioned GUI interaction remains challenging. Current evaluations offer limited coverage, imprecise target-state definitions, and…

Visual Grounding

Composing Parts for Expressive Object Generation

2025-01-01 · CVPR 2025 1 · Harsh Rangwani, Aishwarya Agarwal, Kuldeep Kulkarni, R. Venkatesh Babu 외

Image composition and generation are processes where the artists need control over various parts of the generated images. However, the current state-of-the-art generation models, like Stable Diffusion, cannot handle …

AttributeDenoisingImage GenerationObject

LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models

2026-06-15 · Zhou Tao, Fang Zhang, Zewen Ding, Shida Wang 외 arxiv

Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even when high-resolution inputs preserve the necessary local details. We identify this limitation as visual context rot: deci…

Semantic Jitter: Dense Supervision for Visual Comparisons via Synthetic Images

2016-12-19 · ICCV 2017 10 · Aron Yu, Kristen Grauman

Distinguishing subtle differences in attributes is valuable, yet learning to make visual comparisons remains non-trivial. Not only is the number of possible comparisons quadratic in the number of training images, but als…

AttributeImage GenerationLearning-To-Rank

LoDisc: Learning Global-Local Discriminative Features for Self-Supervised Fine-Grained Visual Recognition

2024-03-06 · Jialu Shi, Zhiqiang Wei, Jie Nie, Lei Huang

Self-supervised contrastive learning strategy has attracted remarkable attention due to its exceptional ability in representation learning. However, current contrastive learning tends to learn global coarse-grained repre…

Contrastive LearningFine-Grained Visual RecognitionObjectObject Recognition+1