VALSE: A Task-Independent Benchmark for Vision and Language Models Centered on Linguistic Phenomena
We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for their visio-linguistic grounding capabilities on specific linguistic phenomena. VALSE offers a suite of six tests covering various linguistic constructs. Solving these requires models to ground linguistic phenomena in the visual modality, allowing more fine-grained evaluations than hitherto possible. We build VALSE using methods that support the construction of valid foils, and report results from evaluating five widely-used V&L models. Our experiments suggest that current models have considerable difficulty addressing most phenomena. Hence, we expect VALSE to serve as an important benchmark to measure future progress of pretrained V&L models from a linguistic perspective, complementing the canonical task-centred V&L evaluations.
Code (1)
Tasks
image-sentence alignmentvalidSimilar Papers 제목 키워드 기반
VALSE: A Task-Independent Benchmark for Vision and Language Models centered on Linguistic Phenomena
We propose VALSE (Vision And Language Structured Evaluation), a novel benchmark designed for testing general-purpose pretrained vision and language (V&L) models for specific visio-linguistic grounding capabilities. Curre…
Question AnsweringVisual Question AnsweringVisual Question Answering (VQA)Visual ReasoningHow and where does CLIP process negation?
Various benchmarks have been proposed to test linguistic understanding in pre-trained vision \& language (VL) models. Here we build on the existence task from the VALSE benchmark (Parcalabescu et al, 2022) which we use t…
Language ModellingNegationDo Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?
Vision and language model (VLM) decoders are currently the best-performing architectures on multimodal tasks. Next to answers, they are able to produce natural language explanations, either in post-hoc or CoT settings. H…
Answer GenerationBenchmarkingExplanation GenerationLanguage ModellingEvalSense: A Framework for Domain-Specific LLM (Meta-)Evaluation
Robust and comprehensive evaluation of large language models (LLMs) is essential for identifying effective LLM system configurations and mitigating risks associated with deploying LLMs in sensitive domains. However, trad…
A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks
We investigate 17 benchmarks (e.g. SugarCREPE, VALSE) commonly used for measuring compositional understanding capabilities of vision-language models (VLMs). We scrutinize design choices in their construction, including d…
Language ModelingLanguage Modelling