paper-with-me

Visual Entailment

3개 벤치마크 · 논문 61편 · 이 태스크의 논문 보기 →

Benchmarks

SNLI-VE val

결과 9개

SNLI-VE test

결과 8개

e-SNLI-VE

결과 2개

Most implemented

Visual Spatial Reasoning

2022-04-30 · 구현 4개

Papers

CogRad: A Cognitively-Inspired Multi-Agent Framework for Radiology Report Generation

2026-07-04 · Saif Ur Rehman Khan, Hasaan Maqsood, Sebastian Vollmer, Andreas Dengel 외 arxiv

Automated radiology report generation (RRG) can ease radiologist workload, yet most existing systems produce a report in a single forward pass, with no mechanism to check a claim against the image or revisit a finding on…

Visual EntailmentVisual Grounding

An Online Reference-Free Evaluation Framework for Flowchart Image-to-Code Generation

2026-02-13 · Giang Son Nguyen, Zi Pong Lim, Sarthak Ketanbhai Modi, Yon Shin Teo 외 arxiv

Vision-Language Models (VLMs) are increasingly used in document processing pipelines to convert flowchart images into structured code (e.g., Mermaid). In production, these systems process arbitrary inputs for which no gr…

Visual EntailmentCode Generation

Dataset Creation for Visual Entailment using Generative AI

2025-08-15 · Rob Reijtenbach, Suzan Verberne, Gijs Wijnholds arxiv

In this paper we present and validate a new synthetic dataset for training visual entailment models. Existing datasets for visual entailment are small and sparse compared to datasets for textual entailment. Manually crea…

Visual Entailment

Probing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls

2025-07-23 · Elena Pitta, Tom Kouwenhoven, Tessa Verhoef arxiv

This study investigates the extent to which the Visual Entailment (VE) task serves as a reliable probe of vision-language understanding in multimodal language models, using the LLaMA 3.2 11B Vision model as a test case. …

Visual EntailmentVisual Grounding

COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark

2025-07-17 · Ishant Chintapatla, Kazuma Choji, Naaisha Agarwal, Andrew Lin 외 arxiv

Recently, many benchmarks and datasets have been developed to evaluate Vision-Language Models (VLMs) using visual question answering (VQA) pairs, and models have shown significant accuracy improvements. However, these be…

Visual Question AnsweringVisual Entailment

Defeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven Optimization

2024-12-19 · Yue Zhang, Liqiang Jing, Vibhav Gogate

We introduce a new task called Defeasible Visual Entailment (DVE), where the goal is to allow the modification of the entailment relationship between an image premise and a text hypothesis based on an additional update. …

Contrastive LearningDecision MakingNatural Language InferenceQuestion Answering+2

전체 61편 보기 →