Papers Referring expression generation
“Referring expression generation” 태그가 달린 논문 86편 · 필터 해제
GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation
Referring Expression Segmentation (RES) and Comprehension (REC) respectively segment and detect the object described by an expression, while Referring Expression Generation (REG) generates an expression for the selected …
Generalized Referring Expression SegmentationReferring expression generationColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation
Referring Expression Comprehension (REC) and Referring Expression Generation (REG) are fundamental tasks in multimodal understanding, supporting precise object localization through natural language. However, existing REC…
Referring expression generationMultimodal ReasoningObject LocalizationVision-Language Models Are Not Pragmatically Competent in Referring Expression Generation
Referring Expression Generation (REG) is a core task for evaluating the pragmatic competence of vision-language systems, requiring not only accurate semantic grounding but also adherence to principles of cooperative comm…
Referring ExpressionReferring expression generationFrontiers in Intelligent Colonoscopy
Colonoscopy is currently one of the most sensitive screening methods for colorectal cancer. This study investigates the frontiers of intelligent colonoscopy techniques and their prospective implications for multimodal me…
Image CaptioningImage ClassificationLanguage Modeling+3Grounding Language in Multi-Perspective Referential Communication
We introduce a task and dataset for referring expression generation and comprehension in multi-agent embodied environments. In this task, two agents in a shared scene must take into account one another's visual perspecti…
Referring ExpressionReferring expression generationUni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoE
Multi-modal large language models (MLLMs) have shown impressive capabilities as a general-purpose interface for various visual and linguistic tasks. However, building a unified MLLM for multi-task learning in the medical…
image-classificationImage ClassificationMixture-of-ExpertsMulti-Task Learning+5Referring Expression Generation in Visually Grounded Dialogue with Discourse-aware Comprehension Guiding
We propose an approach to referring expression generation (REG) in visually grounded dialogue that is meant to produce referring expressions (REs) that are both discriminative and discourse-appropriate. Our method consti…
Image RetrievalReferring ExpressionReferring expression generationRetrievalResilience through Scene Context in Visual Referring Expression Generation
Scene context is well known to facilitate humans' perception of visible objects. In this paper, we investigate the role of context in Referring Expression Generation (REG) for objects in images, where existing research h…
Referring ExpressionReferring expression generationMini-Gemini: Mining the Potential of Multi-modality Vision Language Models
In this work, we introduce Mini-Gemini, a simple and effective framework enhancing multi-modality Vision Language Models (VLMs). Despite the advancements in VLMs facilitating basic visual dialog and reasoning, a performa…
Image ClassificationImage ComprehensionReferring Expression ComprehensionReferring expression generation+2Elysium: Exploring Object-level Perception in Videos via MLLM
Multi-modal Large Language Models (MLLMs) have demonstrated their ability to perceive objects in still images, but their application in video-related tasks, such as object tracking, remains understudied. This lack of exp…
ObjectObject TrackingReferring ExpressionReferring Expression Comprehension+5Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
Multimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs with visual perceiving and grounding capabi…
Language ModelingLanguage ModellingLarge Language ModelMultimodal Large Language Model+2Efficient Multimodal Learning from Data-centric Perspective
Multimodal Large Language Models (MLLMs) have demonstrated notable capabilities in general visual understanding and reasoning tasks. However, their deployment is hindered by substantial computational costs in both traini…
Image ClassificationReferring Expression ComprehensionReferring expression generationIntrinsic Task-based Evaluation for Referring Expression Generation
Recently, a human evaluation study of Referring Expression Generation (REG) models had an unexpected conclusion: on \textsc{webnlg}, Referring Expressions (REs) generated by the state-of-the-art neural models were not on…
Referring ExpressionReferring expression generationText GenerationMobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
We present MobileVLM, a competent multimodal vision language model (MMVLM) targeted to run on mobile devices. It is an amalgamation of a myriad of architectural designs and techniques that are mobile-oriented, which comp…
AutoMLCPUGPUImage Classification+4Enhancing Visual Grounding and Generalization: A Multi-Task Cycle Training Approach for Vision-Language Models
Visual grounding (VG) occupies a pivotal position in multi-modality vision-language models. In this study, we propose ViLaM, a large multi-modality model, that supports multi-tasks of VG using the cycle training strategy…
Image SegmentationLanguage ModellingLarge Language ModelReferring Expression+5GLaMM: Pixel Grounding Large Multimodal Model
Large Multimodal Models (LMMs) extend Large Language Models to the vision domain. Initial LMMs used holistic images and text prompts to generate ungrounded textual responses. Recently, region-level LMMs have been used to…
Conversational Question AnsweringImage CaptioningmodelReferring Expression+4MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Large language models have shown their remarkable capabilities as a general interface for various language-related applications. Motivated by this, we target to build a unified interface for completing many vision-langua…
Image ClassificationImage DescriptionLanguage ModelingLanguage Modelling+8Improved Baselines with Visual Instruction Tuning
Large multimodal models (LMM) have recently shown encouraging progress with visual instruction tuning. In this note, we show that the fully-connected vision-language cross-modal connector in LLaVA is surprisingly powerfu…
Factual Inconsistency Detection in Chart CaptioningImage ClassificationReferring Expression ComprehensionReferring expression generation+4Collecting Visually-Grounded Dialogue with A Game Of Sorts
An idealized, though simplistic, view of the referring expression production and grounding process in (situated) dialogue assumes that a speaker must merely appropriately specify their expression so that the target refer…
Coreference ResolutionImage RetrievalReferring ExpressionReferring Expression Comprehension+4Whether you can locate or not? Interactive Referring Expression Generation
Referring Expression Generation (REG) aims to generate unambiguous Referring Expressions (REs) for objects in a visual scene, with a dual task of Referring Expression Comprehension (REC) to locate the referred object. Ex…
Referring ExpressionReferring Expression ComprehensionReferring expression generation