paper-with-me

Papers Referring Expression Comprehension

“Referring Expression Comprehension” 태그가 달린 논문 167편 · 필터 해제

Referring Expression Instance Retrieval and A Strong End-to-End Baseline

2025-06-23 · Xiangzhao Hao, Kuan Zhu, Hongyu Guo, Haiyun Guo 외

Using natural language to query visual information is a fundamental need in real-world applications. Text-Image Retrieval (TIR) retrieves a target image from a gallery based on an image-level description, while Referring…

Image RetrievalReferring ExpressionReferring Expression ComprehensionRetrieval

Synthetic Visual Genome

2025-06-09 · CVPR 2025 1 · Jae Sung Park, Zixian Ma, Linjie Li, Chenhao Zheng 외

Reasoning over visual relationships-spatial, functional, interactional, social, etc.-is considered to be a fundamental component of human cognition. Yet, despite the major advances in visual comprehension in multimodal l…

Referring ExpressionReferring Expression ComprehensionVisual Reasoning

TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models

2025-05-29 · Yao Xiao, Qiqian Fu, Heyi Tao, Yuqun Wu 외

Image-text models excel at image-level tasks but struggle with detailed visual understanding. While these models provide strong visual-language alignment, segmentation models like SAM2 offer precise spatial boundaries fo…

Referring ExpressionReferring Expression ComprehensionSemantic SegmentationUnsupervised Semantic Segmentation with Language-image Pre-training

WeakMCN: Multi-task Collaborative Network for Weakly Supervised Referring Expression Comprehension and Segmentation

2025-05-24 · CVPR 2025 1 · Yang Liu, Silin Cheng, Xinwei He, Sebastien Ourselin 외

Weakly supervised referring expression comprehension(WREC) and segmentation(WRES) aim to learn object grounding based on a given expression using weak supervision signals like image-text pairs. While these tasks have tra…

Contrastive LearningReferring ExpressionReferring Expression Comprehension

VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

2025-04-10 · Haozhan Shen, Peng Liu, Jingcheng Li, Chunxin Fang 외

Recently DeepSeek R1 has shown that reinforcement learning (RL) can substantially improve the reasoning capabilities of Large Language Models (LLMs) through a simple yet effective design. The core of R1 lies in its rule-…

Language ModelingLanguage ModellingOpen Vocabulary Object DetectionReferring Expression Comprehension+5

Beyond Object Categories: Multi-Attribute Reference Understanding for Visual Grounding

2025-03-25 · Hao Guo, Jianfei Zhu, Wei Fan, Chunzhi Yi 외

Referring expression comprehension (REC) aims at achieving object localization based on natural language descriptions. However, existing REC approaches are constrained by object category descriptions and single-attribute…

AttributeObjectObject LocalizationReferring Expression+2

GeoRSMLLM: A Multimodal Large Language Model for Vision-Language Tasks in Geoscience and Remote Sensing

2025-03-16 · Zilun Zhang, Haozhan Shen, Tiancheng Zhao, Bin Chen 외

The application of Vision-Language Models (VLMs) in remote sensing (RS) has demonstrated significant potential in traditional tasks such as scene classification, object detection, and image captioning. However, current m…

Change DetectionImage CaptioningLanguage ModelingLanguage Modelling+10

New Dataset and Methods for Fine-Grained Compositional Referring Expression Comprehension via Specialist-MLLM Collaboration

2025-02-27 · Xuzheng Yang, Junzhuo Liu, Peng Wang, Guoqing Wang 외

Referring Expression Comprehension (REC) is a foundational cross-modal task that evaluates the interplay of language understanding, image comprehension, and language-to-image grounding. To advance this field, we introduc…

Image ComprehensionReferring ExpressionReferring Expression Comprehension

Exploring Spatial Language Grounding Through Referring Expressions

2025-02-04 · Akshar Tumu, Parisa Kordjamshidi

Spatial Reasoning is an important component of human cognition and is an area in which the latest Vision-language models (VLMs) show signs of difficulty. The current analysis works use image captioning tasks and visual q…

Image CaptioningNegationobject-detectionObject Detection+6

RefDrone: A Challenging Benchmark for Referring Expression Comprehension in Drone Scenes

2025-02-01 · Zhichao Sun, Yepeng Liu, Huachao Zhu, Yuliang Gu 외

Drones have become prevalent robotic platforms with diverse applications, showing significant potential in Embodied Artificial Intelligence (Embodied AI). Referring Expression Comprehension (REC) enables drones to locate…

Referring ExpressionReferring Expression Comprehension

FLORA: Formal Language Model Enables Robust Training-free Zero-shot Object Referring Analysis

2025-01-17 · Zhe Chen, Zijing Chen

Object Referring Analysis (ORA), commonly known as referring expression comprehension, requires the identification and localization of specific objects in an image based on natural descriptions. Unlike generic object det…

Bayesian InferenceLanguage ModelingLanguage ModellingObject+4

Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks

2025-01-14 · CVPR 2025 1 · Miran Heo, Min-Hung Chen, De-An Huang, Sifei Liu 외

We present Omni-RGPT, a multimodal large language model designed to facilitate region-level comprehension for both images and videos. To achieve consistent region representation across spatio-temporal dimensions, we intr…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+3

Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints

2025-01-12 · Ming Dai, Jian Li, Jiedong Zhuang, Xian Zhang 외

Multi-task visual grounding involves the simultaneous execution of localization and segmentation in images based on textual expressions. The majority of advanced methods predominantly focus on transformer-based multimoda…

Image SegmentationReferring ExpressionReferring Expression ComprehensionReferring Expression Segmentation+2

Hierarchical Alignment-enhanced Adaptive Grounding Network for Generalized Referring Expression Comprehension

2025-01-02 · Yaxian Wang, Henghui Ding, Shuting He, Xudong Jiang 외

In this work, we address the challenging task of Generalized Referring Expression Comprehension (GREC). Compared to the classic Referring Expression Comprehension (REC) that focuses on single-target expressions, GREC ext…

Generalized Referring Expression ComprehensionGeneralized Referring Expression SegmentationObject CountingPhrase Grounding+3

DViN: Dynamic Visual Routing Network for Weakly Supervised Referring Expression Comprehension

2025-01-01 · CVPR 2025 1 · Xiaofu Chen, Yaxin Luo, Gen Luo, Jiayi Ji 외

In this paper, we focus on weakly supervised referring expression comprehension (REC), and identify that the lack of fine-grained visual capability greatly limits the upper performance bound of existing methods. To a…

DescriptiveReferring ExpressionReferring Expression ComprehensionReferring Expression Segmentation+1

Task-aware Cross-modal Feature Refinement Transformer with Large Language Models for Visual Grounding

2025-01-01 · CVPR 2025 1 · Wenbo Chen, Zhen Xu, Ruotao Xu, Si Wu 외

The goal of visual grounding is to establish connections between target objects and textual descriptions. Large Language Models (LLMs) have demonstrated strong comprehension abilities across a variety of visual tasks…

Referring ExpressionReferring Expression ComprehensionReferring Expression SegmentationVisual Grounding

Towards Visual Grounding: A Survey

2024-12-28 · Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan, YaoWei Wang 외

Visual Grounding is also known as Referring Expression Comprehension and Phrase Grounding. It involves localizing a natural number of specific regions within an image based on a given textual description. The objective o…

Phrase GroundingReferring ExpressionReferring Expression ComprehensionSurvey+1

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

2024-12-13 · Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu 외

We present DeepSeek-VL2, an advanced series of large Mixture-of-Experts (MoE) Vision-Language Models that significantly improves upon its predecessor, DeepSeek-VL, through two key major upgrades. For the vision component…

Chart UnderstandingMixture-of-ExpertsOptical Character RecognitionQuestion Answering+3

Harlequin: Color-driven Generation of Synthetic Data for Referring Expression Comprehension

2024-11-22 · Luca Parolari, Elena Izzo, Lamberto Ballan

Referring Expression Comprehension (REC) aims to identify a particular object in a scene by a natural language expression, and is an important topic in visual language understanding. State-of-the-art methods for this tas…

Referring ExpressionReferring Expression Comprehension

Frontiers in Intelligent Colonoscopy

2024-10-22 · Ge-Peng Ji, Jingyi Liu, Peng Xu, Nick Barnes 외

Colonoscopy is currently one of the most sensitive screening methods for colorectal cancer. This study investigates the frontiers of intelligent colonoscopy techniques and their prospective implications for multimodal me…

Image CaptioningImage ClassificationLanguage Modeling+3
1–20 / 167 다음 →