Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
The lack of reasoning capabilities in Vision-Language Models (VLMs) has remained at the forefront of research discourse. We posit that this behavior stems from a reporting bias in their training data. That is, how people communicate about visual content by default omits tacit information needed to supervise some types of reasoning; e.g., "at the game today!" is a more likely caption than "a photo of 37 people standing behind a field". We investigate the data underlying the popular VLMs OpenCLIP, LLaVA-1.5 and Molmo through the lens of theories from pragmatics, and find that reporting bias results in insufficient representation of four reasoning skills (spatial, temporal, negation, and counting), despite the corpora being of web-scale, and/or synthetically generated. With a set of curated benchmarks, we demonstrate that: (i) VLMs perform poorly on the aforementioned types of reasoning suppressed in the training data by reporting bias; (ii) contrary to popular belief, scaling data size, model size, and to multiple languages does not result in emergence of these skills by default; but, promisingly, (iii) incorporating annotations specifically collected to obtain tacit information is effective. Our findings highlight the need for more intentional training data curation methods, rather than counting on scale for emergence of reasoning capabilities.
Code (0)
등록된 구현이 없습니다.
Similar Papers 제목 키워드 기반
Do ever larger octopi still amplify reporting biases? Evidence from judgments of typical colour
Language models (LMs) trained on raw texts have no direct access to the physical world. Gordon and Van Durme (2013) point out that LMs can thus suffer from reporting bias: texts rarely report on common facts, instead foc…
Common Sense ReasoningPhysical Commonsense ReasoningProbing Language ModelsDo Neural Language Models Overcome Reporting Bias?
Mining commonsense knowledge from corpora suffers from reporting bias, over-representing the rare at the expense of the trivial (Gordon and Van Durme, 2013). We study to what extent pre-trained language models overcome t…
The Impact of Differential Feature Under-reporting on Algorithmic Fairness
Predictive risk models in the public sector are commonly developed using administrative data that is more complete for subpopulations that more greatly rely on public services. In the United States, for instance, informa…
Decision MakingFairnessChronic pain patient narratives allow for the estimation of current pain intensity
Chronic pain is a multi-dimensional experience, and pain intensity plays an important part, impacting the patients emotional balance, psychology, and behaviour. Standard self-reporting tools, such as the Visual Analogue …
ManagementOn the Extent, Correlates, and Consequences of Reporting Bias in Survey Wages
Surveys are an indispensable source of data for applied economic research; however, their reliance on self-reported information can introduce bias, especially if core variables such as personal income are misreported. To…
Survey