paper-with-me

홈 › Papers

CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space

2026-04-13 · Sohwi Lim, Lee Hyoseok, Jungjoon Park, Tae-Hyun Oh arxiv

Human perception of visual similarity is inherently adaptive and subjective, depending on the users' interests and focus. However, most image retrieval systems fail to reflect this flexibility, relying on a fixed, monolithic metric that cannot incorporate multiple conditions simultaneously. To address this, we propose CLAY, an adaptive similarity computation method that reframes the embedding space of pretrained Vision-Language Models (VLMs) as a text-conditional similarity space without additional training. This design separates the textual conditioning process and visual feature extraction, allowing highly efficient and multi-conditioned retrieval with fixed visual embeddings. We also construct a synthetic evaluation dataset CLAY-EVAL, for comprehensive assessment under diverse conditioned retrieval settings. Experiments on standard datasets and our proposed dataset show that CLAY achieves high retrieval accuracy and notable computational efficiency compared to previous works.

📄 PDF Abstract BibTeX arXiv:2604.11539

Code (0)

등록된 구현이 없습니다.

Tasks

Computational EfficiencyImage Retrieval

Similar Papers 제목 키워드 기반

Visual Sculpting: Visually-Aligned Planning Representations for Long-Horizon Robot Clay Sculpting

2026-05-17 · Peter Schaldenbrand, Jean Oh arxiv

Clay sculpting is a nuanced, artistic task involving dexterous manipulation with long-horizon planning to achieve high-level goals. As a robotics problem, we formulate clay sculpting as a shape-to-shape matching challeng…

Point Clouds

Zero-Shot Satellite Image Retrieval through Joint Embeddings: Application to Crisis Response

2026-05-06 · James Walsh, William Fawcett, Grace Colverd, Raúl Ramos-Pollán arxiv

Semantic search of Earth observation archives remains challenging. Visual foundation models such as CLAY produce rich embeddings of satellite imagery but lack the natural-language grounding needed for intuitive query, an…

Image Retrieval

Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction

2025-11-26 · Gayoung Lee, Junho Kim, Jin-Hwa Kim, Junmo Kim arxiv

Understanding reflection remains a long-standing challenge in 3D reconstruction due to the entanglement of appearance and geometry under view-dependent reflections. In this work, we present the Pygmalion Effect in Vision…

3D Reconstruction

StarGAN-VC2: Rethinking Conditional Methods for StarGAN-Based Voice Conversion

2019-07-29 · Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Nobukatsu Hojo

Non-parallel multi-domain voice conversion (VC) is a technique for learning mappings among multiple domains without relying on parallel data. This is important but challenging owing to the requirement of learning multipl…

Voice Conversion

DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

2025-01-01 · CVPR 2025 1 · Wenhui Liao, Jiapeng Wang, Hongliang Li, Chengyu Wang 외

Text-rich document understanding (TDU) requires comprehensive analysis of documents containing substantial textual content and complex layouts. While Multimodal Large Language Models (MLLMs) have achieved fast progre…

document understandingOptical Character Recognition (OCR)