paper-with-me

홈 › Papers

GOAL: Global-local Object Alignment Learning

2025-03-22 · CVPR 2025 1 · Hyungyu Choi, Young Kyun Jang, Chanho Eom

Vision-language models like CLIP have shown impressive capabilities in aligning images and text, but they often struggle with lengthy and detailed text descriptions because of their training focus on short and concise captions. We present GOAL (Global-local Object Alignment Learning), a novel fine-tuning method that enhances CLIP's ability to handle lengthy text by leveraging both global and local semantic alignments between image and lengthy text. Our approach consists of two key components: Local Image-Sentence Matching (LISM), which identifies corresponding pairs between image segments and descriptive sentences, and Token Similarity-based Learning (TSL), which efficiently propagates local element attention through these matched pairs. Evaluating GOAL on three new benchmarks for image-lengthy text retrieval, we demonstrate significant improvements over baseline CLIP fine-tuning, establishing a simple yet effective approach for adapting CLIP to detailed textual descriptions. Through extensive experiments, we show that our method's focus on local semantic alignment alongside global context leads to more nuanced and representative embeddings, particularly beneficial for tasks requiring fine-grained understanding of lengthy text descriptions.

📄 PDF Abstract BibTeX arXiv:2503.17782

Code (1)

perceptualai-lab/goal 공식 구현 pytorch

Tasks

DescriptiveObjectSentenceText Retrieval

Methods 이 논문이 사용한 방법론

Softmax The Softmax output function transforms a previous layer's output into a vector of probabilities. It is commonly used for multiclass classification. Given an input vector $x$…
Attention 설명 없음
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…
Focus 설명 없음

Similar Papers 제목 키워드 기반

FAST-GOAL: Fast and Efficient Global-local Object Alignment Learning

2026-05-26 · Hyungyu Choi, Young Kyun Jang, Chanho Eom arxiv

Vision-language models such as CLIP have shown impressive capabilities in aligning images and text, but they often struggle with lengthy and detailed text descriptions due to pre-training on short and concise captions. W…

Computational EfficiencyObject Detection

MG-Nav: Dual-Scale Visual Navigation via Sparse Spatial Memory

2025-11-27 · Bo Wang, Jiehong Lin, Chenzhi Liu, Xinting Hu 외 arxiv

We present MG-Nav (Memory-Guided Navigation), a dual-scale framework for zero-shot visual navigation that unifies global memory-guided planning with local geometry-enhanced control. At its core is the Sparse Spatial Memo…

Visual Navigation

DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object Detection

2026-03-19 · Haochen Li, Rui Zhang, Hantao Yao, Xin Zhang 외 arxiv

Domain Adaptive Object Detection (DAOD) aims to transfer detectors from a labeled source domain to an unlabeled target domain. Existing DAOD methods employ multi-granularity feature alignment to learn domain-invariant re…

Long-range modelingObject Detection

Contour Flow: Middle-Level Motion Estimation by Combining Motion Segmentation and Contour Alignment

2015-12-01 · ICCV 2015 12 · Huijun Di, Qingxuan Shi, Feng Lv, Ming Qin 외

Our goal is to estimate contour flow (the contour pairs with consistent point correspondence) from inconsistent contours extracted independently in two video frames. We formulate the contour flow estimation locally as a …

Motion EstimationMotion SegmentationOptical Flow Estimation

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment

2025-05-02 · CVPR 2025 1 · Edson Araujo, Andrew Rouditchenko, Yuan Gong, Saurabhchand Bhati 외

Recent advances in audio-visual learning have shown promising results in learning representations across modalities. However, most approaches rely on global audio representations that fail to capture fine-grained tempora…

audio-visual learningcross-modal alignment