paper-with-me

홈 › Papers

SLAN: Self-Locator Aided Network for Cross-Modal Understanding

2022-11-28 · Jiang-Tian Zhai, Qi Zhang, Tong Wu, Xing-Yu Chen, Jiang-Jiang Liu, Bo Ren, Ming-Ming Cheng

Learning fine-grained interplay between vision and language allows to a more accurate understanding for VisionLanguage tasks. However, it remains challenging to extract key image regions according to the texts for semantic alignments. Most existing works are either limited by textagnostic and redundant regions obtained with the frozen detectors, or failing to scale further due to its heavy reliance on scarce grounding (gold) data to pre-train detectors. To solve these problems, we propose Self-Locator Aided Network (SLAN) for cross-modal understanding tasks without any extra gold data. SLAN consists of a region filter and a region adaptor to localize regions of interest conditioned on different texts. By aggregating cross-modal information, the region filter selects key regions and the region adaptor updates their coordinates with text guidance. With detailed region-word alignments, SLAN can be easily generalized to many downstream tasks. It achieves fairly competitive results on five cross-modal understanding tasks (e.g., 85.7% and 69.2% on COCO image-to-text and text-to-image retrieval, surpassing previous SOTA methods). SLAN also demonstrates strong zero-shot and fine-tuned transferability to two localization tasks.

📄 PDF Abstract BibTeX arXiv:2211.16208

Code (0)

등록된 구현이 없습니다.

Tasks

Image RetrievalImage to textRetrieval

Similar Papers 제목 키워드 기반

SLAN: Self-Locator Aided Network for Vision-Language Understanding

2023-01-01 · ICCV 2023 1 · Jiang-Tian Zhai, Qi Zhang, Tong Wu, Xing-Yu Chen 외

Learning fine-grained interplay between vision and language contributes to a more accurate understanding for Vision-Language tasks. However, it remains challenging to extract key image regions according to the texts …

Image RetrievalImage to textRetrieval

MM-WLAuslan: Multi-View Multi-Modal Word-Level Australian Sign Language Recognition Dataset

2024-10-25 · Xin Shen, Heming Du, Hongwei Sheng, Shuyun Wang 외

Isolated Sign Language Recognition (ISLR) focuses on identifying individual sign language glosses. Considering the diversity of sign languages across geographical regions, developing region-specific ISLR datasets is cruc…

Sign Language Recognition

GeoLocator: a location-integrated large multimodal model for inferring geo-privacy

2023-11-21 · Yifan Yang, Siqin Wang, Daoyang Li, Yixian Zhang 외

Geographic privacy or geo-privacy refers to the keeping private of one's geographic location, especially the restriction of geographical data maintained by personal electronic devices. Geo-privacy is a crucial aspect of …

Image Comprehension

Generalist Scanner Meets Specialist Locator: A Synergistic Coarse-to-Fine Framework for Robust GUI Grounding

2025-09-29 · Zhecheng Li, Guoxian Song, Yiwei Wang, Zhen Xiong 외 arxiv

Grounding natural language queries in graphical user interfaces (GUIs) presents a challenging task that requires models to comprehend diverse UI elements across various applications and systems, while also accurately pre…

Natural Language Queries

AI Pangaea: Unifying Intelligence Islands for Adapting Myriad Tasks

2025-09-22 · Jianlong Chang, Haixin Wang, Zhiyuan Dang, Li Huang 외 arxiv

The pursuit of artificial general intelligence continuously demands generalization in one model across myriad tasks, even those not seen before. However, current AI models are isolated from each other for being limited t…