paper-with-me

홈 › Papers

VR-RAG: Open-vocabulary Species Recognition with RAG-Assisted Large Multi-Modal Models

2025-05-08 · Faizan Farooq Khan, Jun Chen, Youssef Mohamed, Chun-Mei Feng, Mohamed Elhoseiny

Open-vocabulary recognition remains a challenging problem in computer vision, as it requires identifying objects from an unbounded set of categories. This is particularly relevant in nature, where new species are discovered every year. In this work, we focus on open-vocabulary bird species recognition, where the goal is to classify species based on their descriptions without being constrained to a predefined set of taxonomic categories. Traditional benchmarks like CUB-200-2011 and Birdsnap have been evaluated in a closed-vocabulary paradigm, limiting their applicability to real-world scenarios where novel species continually emerge. We show that the performance of current systems when evaluated under settings closely aligned with open-vocabulary drops by a huge margin. To address this gap, we propose a scalable framework integrating structured textual knowledge from Wikipedia articles of 11,202 bird species distilled via GPT-4o into concise, discriminative summaries. We propose Visual Re-ranking Retrieval-Augmented Generation(VR-RAG), a novel, retrieval-augmented generation framework that uses visual similarities to rerank the top m candidates retrieved by a set of multimodal vision language encoders. This allows for the recognition of unseen taxa. Extensive experiments across five established classification benchmarks show that our approach is highly effective. By integrating VR-RAG, we improve the average performance of state-of-the-art Large Multi-Modal Model QWEN2.5-VL by 15.4% across five benchmarks. Our approach outperforms conventional VLM-based approaches, which struggle with unseen species. By bridging the gap between encyclopedic knowledge and visual recognition, our work advances open-vocabulary recognition, offering a flexible, scalable solution for biodiversity monitoring and ecological research.

📄 PDF Abstract BibTeX arXiv:2505.05635

Code (0)

등록된 구현이 없습니다.

Tasks

ArticlesRAGRe-RankingRetrievalRetrieval-augmented Generation

Methods 이 논문이 사용한 방법론

SET Dynamic Sparse Training method where weight mask is updated randomly periodically
Focus 설명 없음

Similar Papers 제목 키워드 기반

Open-Set Recognition of Novel Species in Biodiversity Monitoring

2025-03-03 · Yuyan Chen, Nico Lang, B. Christian Schmidt, Aditya Jain 외

Machine learning is increasingly being applied to facilitate long-term, large-scale biodiversity monitoring. With most species on Earth still undiscovered or poorly documented, species-recognition models are expected to …

Fine-Grained Image RecognitionOpen Set LearningOut-of-Distribution Detection

ORCA: Object Recognition and Comprehension for Archiving Marine Species

2025-12-24 · Yuk-Kwan Wong, Haixin Liang, Zeyu Ma, Yiwei Chen 외 arxiv

Marine visual understanding is essential for monitoring and protecting marine ecosystems, enabling automatic and scalable biological surveys. However, progress is hindered by limited training data and the lack of a syste…

Object RecognitionObject DetectionVisual Grounding

Open-Vocabulary Segmentation with Semantic-Assisted Calibration

2023-12-07 · CVPR 2024 1 · Yong liu, Sule Bai, Guanbin Li, Yitong Wang 외

This paper studies open-vocabulary segmentation (OVS) through calibrating in-vocabulary and domain-biased embedding space with generalized contextual prior of CLIP. As the core of open-vocabulary understanding, alignment…

AttributeOpen Vocabulary Semantic Segmentation

Detection, Recognition and Tracking of Moving Objects from Real-time Video via Visual Vocabulary Model and Species Inspired PSO

2017-06-02 · Kumar S. Ray, Anit Chakraborty, Sayandip Dutta

In this paper, we address the basic problem of recognizing moving objects in video images using Visual Vocabulary model and Bag of Words and track our object of interest in the subsequent video frames using species inspi…

ClassificationGeneral ClassificationObject

Open-Vocabulary Animal Keypoint Detection with Semantic-feature Matching

2023-10-08 · Hao Zhang, Lumin Xu, Shenqi Lai, Wenqi Shao 외

Current image-based keypoint detection methods for animal (including human) bodies and faces are generally divided into full-supervised and few-shot class-agnostic approaches. The former typically relies on laborious and…

Keypoint DetectionOpen Vocabulary Keypoint Detection