paper-with-me

SME

Standard Multimodal Explanation

홈페이지 · 논문 7편

SME is a new dataset for Multi-modal Explanation for Visual Question Answering comprising 1,028,230 samples, with 1,656 visual objects requiring detection in explanations. To our knowledge, this is the first dataset where the explanations are in standard English with additional visual grounding tokens.

ImagesTexts English

벤치마크

FS-MEVQA on SME 결과 21개