paper-with-me

Papers Referring expression generation

“Referring expression generation” 태그가 달린 논문 86편 · 필터 해제

GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation

2026-01-08 · Henghui Ding, Chang Liu, Shuting He, Xudong Jiang 외 arxiv

Referring Expression Segmentation (RES) and Comprehension (REC) respectively segment and detect the object described by an expression, while Referring Expression Generation (REG) generates an expression for the selected …

Generalized Referring Expression SegmentationReferring expression generation

ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation

2025-09-28 · Shilan Zhang, Jirui Huang, Ruilin Yao, Cong Wang 외 arxiv

Referring Expression Comprehension (REC) and Referring Expression Generation (REG) are fundamental tasks in multimodal understanding, supporting precise object localization through natural language. However, existing REC…

Referring expression generationMultimodal ReasoningObject Localization

Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation

2025-04-22 · Ziqiao Ma, Jing Ding, Xuejun Zhang, Dezhi Luo 외

Referring Expression Generation (REG) is a core task for evaluating the pragmatic competence of vision-language systems, requiring not only accurate semantic grounding but also adherence to principles of cooperative comm…

Referring ExpressionReferring expression generation

Frontiers in Intelligent Colonoscopy

2024-10-22 · Ge-Peng Ji, Jingyi Liu, Peng Xu, Nick Barnes 외

Colonoscopy is currently one of the most sensitive screening methods for colorectal cancer. This study investigates the frontiers of intelligent colonoscopy techniques and their prospective implications for multimodal me…

Image CaptioningImage ClassificationLanguage Modeling+3

Grounding Language in Multi-Perspective Referential Communication

2024-10-04 · Zineng Tang, Lingjun Mao, Alane Suhr

We introduce a task and dataset for referring expression generation and comprehension in multi-agent embodied environments. In this task, two agents in a shared scene must take into account one another's visual perspecti…

Referring ExpressionReferring expression generation

Uni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoE

2024-09-26 · Xun Zhu, Ying Hu, Fanbin Mo, Miao Li 외

Multi-modal large language models (MLLMs) have shown impressive capabilities as a general-purpose interface for various visual and linguistic tasks. However, building a unified MLLM for multi-task learning in the medical…

image-classificationImage ClassificationMixture-of-ExpertsMulti-Task Learning+5

Referring Expression Generation in Visually Grounded Dialogue with Discourse-aware Comprehension Guiding

2024-09-09 · INLG 2024 9 · Bram Willemsen, Gabriel Skantze

We propose an approach to referring expression generation (REG) in visually grounded dialogue that is meant to produce referring expressions (REs) that are both discriminative and discourse-appropriate. Our method consti…

Image RetrievalReferring ExpressionReferring expression generationRetrieval

Resilience through Scene Context in Visual Referring Expression Generation

2024-04-18 · Simeon Junker, Sina Zarrieß

Scene context is well known to facilitate humans' perception of visible objects. In this paper, we investigate the role of context in Referring Expression Generation (REG) for objects in images, where existing research h…

Referring ExpressionReferring expression generation

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

2024-03-27 · Yanwei Li, Yuechen Zhang, Chengyao Wang, Zhisheng Zhong 외

In this work, we introduce Mini-Gemini, a simple and effective framework enhancing multi-modality Vision Language Models (VLMs). Despite the advancements in VLMs facilitating basic visual dialog and reasoning, a performa…

Image ClassificationImage ComprehensionReferring Expression ComprehensionReferring expression generation+2

Elysium: Exploring Object-level Perception in Videos via MLLM

2024-03-25 · Han Wang, Yanjie Wang, YongJie Ye, Yuxiang Nie 외

Multi-modal Large Language Models (MLLMs) have demonstrated their ability to perceive objects in still images, but their application in video-related tasks, such as object tracking, remains understudied. This lack of exp…

ObjectObject TrackingReferring ExpressionReferring Expression Comprehension+5

Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception

2024-03-05 · CVPR 2024 1 · Junwen He, Yifan Wang, Lijun Wang, Huchuan Lu 외

Multimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs with visual perceiving and grounding capabi…

Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+2

Efficient Multimodal Learning from Data-centric Perspective

2024-02-18 · Muyang He, Yexin Liu, Boya Wu, Jianhao Yuan 외

Multimodal Large Language Models (MLLMs) have demonstrated notable capabilities in general visual understanding and reasoning tasks. However, their deployment is hindered by substantial computational costs in both traini…

Image ClassificationReferring Expression ComprehensionReferring expression generation

Intrinsic Task-based Evaluation for Referring Expression Generation

2024-02-12 · Guanyi Chen, Fahime Same, Kees Van Deemter

Recently, a human evaluation study of Referring Expression Generation (REG) models had an unexpected conclusion: on \textsc{webnlg}, Referring Expressions (REs) generated by the state-of-the-art neural models were not on…

Referring ExpressionReferring expression generationText Generation

MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

2023-12-28 · Xiangxiang Chu, Limeng Qiao, Xinyang Lin, Shuang Xu 외

We present MobileVLM, a competent multimodal vision language model (MMVLM) targeted to run on mobile devices. It is an amalgamation of a myriad of architectural designs and techniques that are mobile-oriented, which comp…

AutoMLCPUGPUImage Classification+4

Enhancing Visual Grounding and Generalization: A Multi-Task Cycle Training Approach for Vision-Language Models

2023-11-21 · Xiaoyu Yang, Lijian Xu, Hao Sun, Hongsheng Li 외

Visual grounding (VG) occupies a pivotal position in multi-modality vision-language models. In this study, we propose ViLaM, a large multi-modality model, that supports multi-tasks of VG using the cycle training strategy…

Image SegmentationLanguage ModellingLarge Language ModelReferring Expression+5

GLaMM: Pixel Grounding Large Multimodal Model

2023-11-06 · CVPR 2024 1 · Hanoona Rasheed, Muhammad Maaz, Sahal Shaji Mullappilly, Abdelrahman Shaker 외

Large Multimodal Models (LMMs) extend Large Language Models to the vision domain. Initial LMMs used holistic images and text prompts to generate ungrounded textual responses. Recently, region-level LMMs have been used to…

Conversational Question AnsweringImage CaptioningmodelReferring Expression+4

MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

2023-10-14 · Jun Chen, Deyao Zhu, Xiaoqian Shen, Xiang Li 외

Large language models have shown their remarkable capabilities as a general interface for various language-related applications. Motivated by this, we target to build a unified interface for completing many vision-langua…

Image ClassificationImage DescriptionLanguage ModelingLanguage Modelling+8

Improved Baselines with Visual Instruction Tuning

2023-10-05 · CVPR 2024 1 · Haotian Liu, Chunyuan Li, Yuheng Li, Yong Jae Lee

Large multimodal models (LMM) have recently shown encouraging progress with visual instruction tuning. In this note, we show that the fully-connected vision-language cross-modal connector in LLaVA is surprisingly powerfu…

Factual Inconsistency Detection in Chart CaptioningImage ClassificationReferring Expression ComprehensionReferring expression generation+4

Collecting Visually-Grounded Dialogue with A Game Of Sorts

2023-09-10 · LREC 2022 6 · Bram Willemsen, Dmytro Kalpakchi, Gabriel Skantze

An idealized, though simplistic, view of the referring expression production and grounding process in (situated) dialogue assumes that a speaker must merely appropriately specify their expression so that the target refer…

Coreference ResolutionImage RetrievalReferring ExpressionReferring Expression Comprehension+4

Whether you can locate or not? Interactive Referring Expression Generation

2023-08-19 · Fulong Ye, Yuxing Long, Fangxiang Feng, Xiaojie Wang

Referring Expression Generation (REG) aims to generate unambiguous Referring Expressions (REs) for objects in a visual scene, with a dual task of Referring Expression Comprehension (REC) to locate the referred object. Ex…

Referring ExpressionReferring Expression ComprehensionReferring expression generation
1–20 / 86 다음 →