paper-with-me

홈 › Papers

The BLA Benchmark: Investigating Basic Language Abilities of Pre-Trained Multimodal Models

2023-10-23 · Xinyi Chen, Raquel Fernández, Sandro Pezzelle

Despite the impressive performance achieved by pre-trained language-and-vision models in downstream tasks, it remains an open question whether this reflects a proper understanding of image-text interaction. In this work, we explore to what extent they handle basic linguistic constructions -- active-passive voice, coordination, and relative clauses -- that even preschool children can typically master. We present BLA, a novel, automatically constructed benchmark to evaluate multimodal models on these Basic Language Abilities. We show that different types of Transformer-based systems, such as CLIP, ViLBERT, and BLIP2, generally struggle with BLA in a zero-shot setting, in line with previous findings. Our experiments, in particular, show that most of the tested models only marginally benefit when fine-tuned or prompted with construction-specific samples. Yet, the generative BLIP2 shows promising trends, especially in an in-context learning setting. This opens the door to using BLA not only as an evaluation benchmark but also to improve models' basic language abilities.

📄 PDF Abstract BibTeX arXiv:2310.15061

Code (1)

shin-ee-chen/bla 공식 구현 pytorch

Tasks

In-Context Learning

Methods 이 논문이 사용한 방법론

ViLBERT Vision-and-Language BERT (ViLBERT) is a BERT-based model for learning task-agnostic joint representations of image content and…
CLIP Contrastive Language-Image Pre-training (CLIP), consisting of a simplified version of ConVIRT trained from scratch, is an efficient method of image representation learning…

Similar Papers 제목 키워드 기반

Investigating the (De)Composition Capabilities of Large Language Models in Natural-to-Formal Language Conversion

2025-01-24 · Ziyao Xu, Houfeng Wang

To achieve generalized and robust natural-to-formal language conversion (N2F), large language models (LLMs) need to have strong capabilities of decomposition and composition in N2F when faced with an unfamiliar formal la…

Natural Language Understanding

Response Wide Shut? Surprising Observations in Basic Vision Language Model Capabilities

2025-07-10 · Shivam Chandhok, Wan-Cyuan Fan, Vered Shwartz, Vineeth N Balasubramanian 외 arxiv

Vision-language Models (VLMs) have emerged as general-purpose tools for addressing a variety of complex computer vision problems. Such models have been shown to be highly capable, but, at the same time, lacking some basi…

What's "up" with vision-language models? Investigating their struggle with spatial reasoning

2023-10-30 · Amita Kamath, Jack Hessel, Kai-Wei Chang

Recent vision-language (VL) models are powerful, but can they reliably distinguish "right" from "left"? We curate three new corpora to quantify model comprehension of such basic spatial relations. These tests isolate spa…

Spatial Reasoning

CFBenchmark: Chinese Financial Assistant Benchmark for Large Language Model

2023-11-10 · Yang Lei, Jiangtong Li, Dawei Cheng, Zhijun Ding 외

Large language models (LLMs) have demonstrated great potential in the financial domain. Thus, it becomes important to assess the performance of LLMs in the financial tasks. In this work, we introduce CFBenchmark, to eval…

Language ModelingLanguage ModellingLarge Language Model

Towards a Robust Detection of Language Model Generated Text: Is ChatGPT that Easy to Detect?

2023-06-09 · Wissam Antoun, Virginie Mouilleron, Benoît Sagot, Djamé Seddah

Recent advances in natural language processing (NLP) have led to the development of large language models (LLMs) such as ChatGPT. This paper proposes a methodology for developing and evaluating ChatGPT detectors for Fren…

Adversarial TextLanguage ModelingLanguage Modelling