paper-with-me

Papers Zero-Shot Visual Question Answring

“Zero-Shot Visual Question Answring” 태그가 달린 논문 3편 · 필터 해제

CoLLaVO: Crayon Large Language and Vision mOdel

2024-02-17 · Byung-Kwan Lee, Beomchan Park, Chae Won Kim, Yong Man Ro

The remarkable success of Large Language Models (LLMs) and instruction tuning drives the evolution of Vision Language Models (VLMs) towards a versatile general-purpose model. Yet, it remains unexplored whether current VL…

Large Language ModelmodelObjectVisual Prompt Tuning+3

Implicit Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis

2023-09-21 · NeurIPS 2023 11

Deep network models are often purely inductive during both training and inference on unseen data. When these models are used for prediction, but they may fail to capture important semantic information and implicit depend…

Cross-Modal RetrievalImage CaptioningImage RetrievalNatural Language Understanding+5

MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

2023-08-04 · Weihao Yu, Zhengyuan Yang, Linjie Li, JianFeng Wang 외

We propose MM-Vet, an evaluation benchmark that examines large multimodal models (LMMs) on complicated multimodal tasks. Recent LMMs have shown various intriguing abilities, such as solving math problems written on the b…

MathMM-VetZero-Shot Visual Question Answring
1–3 / 3