paper-with-me

홈 › Papers

SLAN: Self-Locator Aided Network for Vision-Language Understanding

2023-01-01 · ICCV 2023 1 · Jiang-Tian Zhai, Qi Zhang, Tong Wu, Xing-Yu Chen, Jiang-Jiang Liu, Ming-Ming Cheng

Learning fine-grained interplay between vision and language contributes to a more accurate understanding for Vision-Language tasks. However, it remains challenging to extract key image regions according to the texts for semantic alignments. Most existing works are either limited by text-agnostic and redundant regions obtained with the frozen detectors, or failing to scale further due to their heavy reliance on scarce grounding (gold) data to pre-train detectors. To solve these problems, we propose Self-Locator Aided Network (SLAN) for vision-language understanding tasks without any extra gold data. SLAN consists of a region filter and a region adaptor to localize regions of interest conditioned on different texts. By aggregating vision-language information, the region filter selects key regions and the region adaptor updates their coordinates with text guidance. With detailed region-word alignments, SLAN can be easily generalized to many downstream tasks. It achieves fairly competitive results on five vision-language understanding tasks (e.g., 85.7% and 69.2% on COCO image-to-text and text-to-image retrieval, surpassing previous SOTA methods). SLAN also demonstrates strong zero-shot and fine-tuned transferability to two localization tasks.

📄 PDF Abstract BibTeX

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalImage to textRetrieval

Similar Papers 제목 키워드 기반

SLAN: Self-Locator Aided Network for Cross-Modal Understanding

2022-11-28 · Jiang-Tian Zhai, Qi Zhang, Tong Wu, Xing-Yu Chen 외

Learning fine-grained interplay between vision and language allows to a more accurate understanding for VisionLanguage tasks. However, it remains challenging to extract key image regions according to the texts for semant…

Image RetrievalImage to textRetrieval

SC-Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language Models

2024-03-20 · CVPR 2024 1 · Tongtian Yue, Jie Cheng, Longteng Guo, Xingyuan Dai 외

Recent trends in Large Vision Language Models (LVLMs) research have been increasingly focusing on advancing beyond general image understanding towards more nuanced, object-level referential comprehension. In this paper, …

Anchored, Not Graded: Vision-Language Models Fail at Slant-from-Texture Perception

2026-06-04 · Qian Zhang, Michal Golovanevsky, Fulvio Domini, James Tompkin arxiv

Human perception of surface slant from texture exhibits systematic, graded biases that emerge reliably in psychophysical experiments. Prior work showed that unsupervised CNNs reproduce several human-like biases, while su…

Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding

2025-09-29 · Zhecheng Li, Guoxian Song, Yiwei Wang, Zhen Xiong 외 arxiv

Grounding natural language queries in graphical user interfaces (GUIs) presents a challenging task that requires models to comprehend diverse UI elements across various applications and systems, while also accurately pre…

Natural Language Queries

Approximating Trajectory Constraints with Machine Learning -- Microgrid Islanding with Frequency Constraints

2020-01-16 · Yichen Zhang, Chen Chen, Guodong Liu, Tianqi Hong 외

In this paper, we introduce a deep learning aided constraint encoding method to tackle the frequency-constraint microgrid scheduling problem. The nonlinear function between system operating condition and frequency nadir …

BIG-bench Machine LearningScheduling