Referring expression generation
2개 벤치마크 · 논문 86편 · 이 태스크의 논문 보기 →
Benchmarks
ColonINST-v1 (Seen)
ColonINST-v1 (Unseen)
Most implemented
Visual Instruction Tuning
Improved Baselines with Visual Instruction Tuning
Modeling Context in Referring Expressions
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Kosmos-2: Grounding Multimodal Large Language Models to the World
Papers
GREx: Generalized Referring Expression Segmentation, Comprehension, and Generation
Referring Expression Segmentation (RES) and Comprehension (REC) respectively segment and detect the object described by an expression, while Referring Expression Generation (REG) generates an expression for the selected …
Generalized Referring Expression SegmentationReferring expression generationColLab: A Collaborative Spatial Progressive Data Engine for Referring Expression Comprehension and Generation
Referring Expression Comprehension (REC) and Referring Expression Generation (REG) are fundamental tasks in multimodal understanding, supporting precise object localization through natural language. However, existing REC…
Referring expression generationMultimodal ReasoningObject LocalizationVision-Language Models Are Not Pragmatically Competent in Referring Expression Generation
Referring Expression Generation (REG) is a core task for evaluating the pragmatic competence of vision-language systems, requiring not only accurate semantic grounding but also adherence to principles of cooperative comm…
Referring ExpressionReferring expression generationFrontiers in Intelligent Colonoscopy
Colonoscopy is currently one of the most sensitive screening methods for colorectal cancer. This study investigates the frontiers of intelligent colonoscopy techniques and their prospective implications for multimodal me…
Image CaptioningImage ClassificationLanguage Modeling+3Grounding Language in Multi-Perspective Referential Communication
We introduce a task and dataset for referring expression generation and comprehension in multi-agent embodied environments. In this task, two agents in a shared scene must take into account one another's visual perspecti…
Referring ExpressionReferring expression generationUni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoE
Multi-modal large language models (MLLMs) have shown impressive capabilities as a general-purpose interface for various visual and linguistic tasks. However, building a unified MLLM for multi-task learning in the medical…
image-classificationImage ClassificationMixture-of-ExpertsMulti-Task Learning+5