Papers Referring Expression Comprehension
“Referring Expression Comprehension” 태그가 달린 논문 167편 · 필터 해제
Referring Expression Instance Retrieval and A Strong End-to-End Baseline
Using natural language to query visual information is a fundamental need in real-world applications. Text-Image Retrieval (TIR) retrieves a target image from a gallery based on an image-level description, while Referring…
Image RetrievalReferring ExpressionReferring Expression ComprehensionRetrievalSynthetic Visual Genome
Reasoning over visual relationships-spatial, functional, interactional, social, etc.-is considered to be a fundamental component of human cognition. Yet, despite the major advances in visual comprehension in multimodal l…
Referring ExpressionReferring Expression ComprehensionVisual ReasoningTextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
Image-text models excel at image-level tasks but struggle with detailed visual understanding. While these models provide strong visual-language alignment, segmentation models like SAM2 offer precise spatial boundaries fo…
Referring ExpressionReferring Expression ComprehensionSemantic SegmentationUnsupervised Semantic Segmentation with Language-image Pre-trainingWeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation
Weakly supervised referring expression comprehension(WREC) and segmentation(WRES) aim to learn object grounding based on a given expression using weak supervision signals like image-text pairs. While these tasks have tra…
Contrastive LearningReferring ExpressionReferring Expression ComprehensionVLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Recently DeepSeek R1 has shown that reinforcement learning (RL) can substantially improve the reasoning capabilities of Large Language Models (LLMs) through a simple yet effective design. The core of R1 lies in its rule-…
Language ModelingLanguage ModellingOpen Vocabulary Object DetectionReferring Expression Comprehension+5Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding
Referring expression comprehension (REC) aims at achieving object localization based on natural language descriptions. However, existing REC approaches are constrained by object category descriptions and single-attribute…
AttributeObjectObject LocalizationReferring Expression+2GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing
The application of Vision-Language Models (VLMs) in remote sensing (RS) has demonstrated significant potential in traditional tasks such as scene classification, object detection, and image captioning. However, current m…
Change DetectionImage CaptioningLanguage ModelingLanguage Modelling+10New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM Collaboration
Referring Expression Comprehension (REC) is a foundational cross-modal task that evaluates the interplay of language understanding, image comprehension, and language-to-image grounding. To advance this field, we introduc…
Image ComprehensionReferring ExpressionReferring Expression ComprehensionExploring Spatial Language Grounding Through Referring Expressions
Spatial Reasoning is an important component of human cognition and is an area in which the latest Vision-language models (VLMs) show signs of difficulty. The current analysis works use image captioning tasks and visual q…
Image CaptioningNegationobject-detectionObject Detection+6RefDrone: A Challenging Benchmark for Referring Expression Comprehension in Drone Scenes
Drones have become prevalent robotic platforms with diverse applications, showing significant potential in Embodied Artificial Intelligence (Embodied AI). Referring Expression Comprehension (REC) enables drones to locate…
Referring ExpressionReferring Expression ComprehensionFLORA: Formal Language Model Enables Robust Training-free Zero-shot Object Referring Analysis
Object Referring Analysis (ORA), commonly known as referring expression comprehension, requires the identification and localization of specific objects in an image based on natural descriptions. Unlike generic object det…
Bayesian InferenceLanguage ModelingLanguage ModellingObject+4Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks
We present Omni-RGPT, a multimodal large language model designed to facilitate region-level comprehension for both images and videos. To achieve consistent region representation across spatio-temporal dimensions, we intr…
Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+3Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints
Multi-task visual grounding involves the simultaneous execution of localization and segmentation in images based on textual expressions. The majority of advanced methods predominantly focus on transformer-based multimoda…
Image SegmentationReferring ExpressionReferring Expression ComprehensionReferring Expression Segmentation+2Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension
In this work, we address the challenging task of Generalized Referring Expression Comprehension (GREC). Compared to the classic Referring Expression Comprehension (REC) that focuses on single-target expressions, GREC ext…
Generalized Referring Expression ComprehensionGeneralized Referring Expression SegmentationObject CountingPhrase Grounding+3DViN: Dynamic Visual Routing Network for Weakly Supervised Referring Expression Comprehension
In this paper, we focus on weakly supervised referring expression comprehension (REC), and identify that the lack of fine-grained visual capability greatly limits the upper performance bound of existing methods. To a…
DescriptiveReferring ExpressionReferring Expression ComprehensionReferring Expression Segmentation+1Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual Grounding
The goal of visual grounding is to establish connections between target objects and textual descriptions. Large Language Models (LLMs) have demonstrated strong comprehension abilities across a variety of visual tasks…
Referring ExpressionReferring Expression ComprehensionReferring Expression SegmentationVisual GroundingTowards Visual Grounding: A Survey
Visual Grounding is also known as Referring Expression Comprehension and Phrase Grounding. It involves localizing a natural number of specific regions within an image based on a given textual description. The objective o…
Phrase GroundingReferring ExpressionReferring Expression ComprehensionSurvey+1DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
We present DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL, through two key major upgrades. For the vision component…
Chart UnderstandingMixture-of-ExpertsOptical Character RecognitionQuestion Answering+3Harlequin: Color-driven Generation of Synthetic Data for Referring Expression Comprehension
Referring Expression Comprehension (REC) aims to identify a particular object in a scene by a natural language expression, and is an important topic in visual language understanding. State-of-the-art methods for this tas…
Referring ExpressionReferring Expression ComprehensionFrontiers in Intelligent Colonoscopy
Colonoscopy is currently one of the most sensitive screening methods for colorectal cancer. This study investigates the frontiers of intelligent colonoscopy techniques and their prospective implications for multimodal me…
Image CaptioningImage ClassificationLanguage Modeling+3