Visual Entailment
3개 벤치마크 · 논문 61편 · 이 태스크의 논문 보기 →
Benchmarks
Most implemented
UNITER: UNiversal Image-TExt Representation Learning
CoCa: Contrastive Captioners are Image-Text Foundation Models
Visual Spatial Reasoning
OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
Distilled Dual-Encoder Model for Vision-Language Understanding
How Much Can CLIP Benefit Vision-and-Language Tasks?
Papers
CogRad: A Cognitively-Inspired Multi-Agent Framework for Radiology Report Generation
Automated radiology report generation (RRG) can ease radiologist workload, yet most existing systems produce a report in a single forward pass, with no mechanism to check a claim against the image or revisit a finding on…
Visual EntailmentVisual GroundingAn Online Reference-Free Evaluation Framework for Flowchart Image-to-Code Generation
Vision-Language Models (VLMs) are increasingly used in document processing pipelines to convert flowchart images into structured code (e.g., Mermaid). In production, these systems process arbitrary inputs for which no gr…
Visual EntailmentCode GenerationDataset Creation for Visual Entailment using Generative AI
In this paper we present and validate a new synthetic dataset for training visual entailment models. Existing datasets for visual entailment are small and sparse compared to datasets for textual entailment. Manually crea…
Visual EntailmentProbing Vision-Language Understanding through the Visual Entailment Task: promises and pitfalls
This study investigates the extent to which the Visual Entailment (VE) task serves as a reliable probe of vision-language understanding in multimodal language models, using the LLaMA 3.2 11B Vision model as a test case. …
Visual EntailmentVisual GroundingCOREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
Recently, many benchmarks and datasets have been developed to evaluate Vision-Language Models (VLMs) using visual question answering (VQA) pairs, and models have shown significant accuracy improvements. However, these be…
Visual Question AnsweringVisual EntailmentDefeasible Visual Entailment: Benchmark, Evaluator, and Reward-Driven Optimization
We introduce a new task called Defeasible Visual Entailment (DVE), where the goal is to allow the modification of the entailment relationship between an image premise and a text hypothesis based on an additional update. …
Contrastive LearningDecision MakingNatural Language InferenceQuestion Answering+2