paper-with-me

e-ViL

홈페이지 · 논문 10편

e-ViL is a benchmark for explainable vision-language tasks. e-ViL spans across three datasets of human-written NLEs (natural language explanations), and provides a unified evaluation framework that is designed to be re-usable for future works. This benchmark uses the following datasets: [e-SNLI-VE](e-snli-ve), [VCR](vcr), VQA-X.

ImagesTexts English