paper-with-me

Referring expression generation

2개 벤치마크 · 논문 86편 · 이 태스크의 논문 보기 →

Benchmarks

ColonINST-v1 (Seen)

결과 17개

ColonINST-v1 (Unseen)

결과 17개

Most implemented

Visual Instruction Tuning

2023-04-17 · 구현 13개

Modeling Context in Referring Expressions

2016-07-31 · 구현 4개

Papers

GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation

2026-01-08 · Henghui Ding, Chang Liu, Shuting He, Xudong Jiang 외 arxiv

Referring Expression Segmentation (RES) and Comprehension (REC) respectively segment and detect the object described by an expression, while Referring Expression Generation (REG) generates an expression for the selected …

Generalized Referring Expression SegmentationReferring expression generation

ColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation

2025-09-28 · Shilan Zhang, Jirui Huang, Ruilin Yao, Cong Wang 외 arxiv

Referring Expression Comprehension (REC) and Referring Expression Generation (REG) are fundamental tasks in multimodal understanding, supporting precise object localization through natural language. However, existing REC…

Referring expression generationMultimodal ReasoningObject Localization

Vision-Language Models Are Not Pragmatically Competent in Referring Expression Generation

2025-04-22 · Ziqiao Ma, Jing Ding, Xuejun Zhang, Dezhi Luo 외

Referring Expression Generation (REG) is a core task for evaluating the pragmatic competence of vision-language systems, requiring not only accurate semantic grounding but also adherence to principles of cooperative comm…

Referring ExpressionReferring expression generation

Frontiers in Intelligent Colonoscopy

2024-10-22 · Ge-Peng Ji, Jingyi Liu, Peng Xu, Nick Barnes 외

Colonoscopy is currently one of the most sensitive screening methods for colorectal cancer. This study investigates the frontiers of intelligent colonoscopy techniques and their prospective implications for multimodal me…

Image CaptioningImage ClassificationLanguage Modeling+3

Grounding Language in Multi-Perspective Referential Communication

2024-10-04 · Zineng Tang, Lingjun Mao, Alane Suhr

We introduce a task and dataset for referring expression generation and comprehension in multi-agent embodied environments. In this task, two agents in a shared scene must take into account one another's visual perspecti…

Referring ExpressionReferring expression generation

Uni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoE

2024-09-26 · Xun Zhu, Ying Hu, Fanbin Mo, Miao Li 외

Multi-modal large language models (MLLMs) have shown impressive capabilities as a general-purpose interface for various visual and linguistic tasks. However, building a unified MLLM for multi-task learning in the medical…

image-classificationImage ClassificationMixture-of-ExpertsMulti-Task Learning+5

전체 86편 보기 →