Papers Zero-Shot Visual Question Answring
“Zero-Shot Visual Question Answring” 태그가 달린 논문 3편 · 필터 해제
CoLLaVO: Crayon Large Language and Vision mOdel
2024-02-17
· Byung-Kwan Lee, Beomchan Park, Chae Won Kim, Yong Man Ro
The remarkable success of Large Language Models (LLMs) and instruction tuning drives the evolution of Vision Language Models (VLMs) towards a versatile general-purpose model. Yet, it remains unexplored whether current VL…
Large Language ModelmodelObjectVisual Prompt Tuning+3Implicit Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis
2023-09-21 · NeurIPS 2023 11
Deep network models are often purely inductive during both training and inference on unseen data. When these models are used for prediction, but they may fail to capture important semantic information and implicit depend…
Cross-Modal RetrievalImage CaptioningImage RetrievalNatural Language Understanding+5MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
2023-08-04
· Weihao Yu, Zhengyuan Yang, Linjie Li, JianFeng Wang 외
We propose MM-Vet, an evaluation benchmark that examines large multimodal models (LMMs) on complicated multimodal tasks. Recent LMMs have shown various intriguing abilities, such as solving math problems written on the b…
MathMM-VetZero-Shot Visual Question Answring
1–3 / 3