Grounded Intuition of GPT-Vision's Abilities with Scientific Images
GPT-Vision has impressed us on a range of vision-language tasks, but it comes with the familiar new challenge: we have little idea of its capabilities and limitations. In our study, we formalize a process that many have instinctively been trying already to develop "grounded intuition" of this new model. Inspired by the recent movement away from benchmarking in favor of example-driven qualitative evaluation, we draw upon grounded theory and thematic analysis in social science and human-computer interaction to establish a rigorous framework for qualitative evaluation in natural language processing. We use our technique to examine alt text generation for scientific figures, finding that GPT-Vision is particularly sensitive to prompting, counterfactual text in images, and relative spatial relationships. Our method and analysis aim to help researchers ramp up their own grounded intuitions of new models while exposing how GPT-Vision can be applied to make information more accessible.
Code (1)
Tasks
BenchmarkingcounterfactualText GenerationSimilar Papers 제목 키워드 기반
From Macro to Micro: Benchmarking Microscopic Spatial Intelligence on Molecules via Vision-Language Models
This paper introduces the concept of Microscopic Spatial Intelligence (MiSI), the capability to perceive and reason about the spatial relationships of invisible microscopic entities, which is fundamental to scientific di…
SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
Most of the existing Large Language Model (LLM) benchmarks on scientific problem reasoning focus on problems grounded in high-school subjects and are confined to elementary algebraic operations. To systematically examine…
BenchmarkingLanguage ModelingLanguage ModellingLarge Language Model+1SciCustom: A Framework for Custom Evaluation of Scientific Capabilities in Large Language Models
Large language models (LLMs) are increasingly applied to scientific research, yet existing evaluations often fail to reflect the fine-grained capabilities required in practice. Most benchmarks are manually curated or dom…
Question GenerationA Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images
Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competit…
Visual Question AnsweringInformation ExtractionCan Theoretical Physics Research Benefit from Language Agents?
Large Language Models (LLMs) are rapidly advancing across diverse domains, yet their application in theoretical physics research is not yet mature. This position paper argues that LLM agents can potentially help accelera…
Code GenerationMathematical ReasoningPhysical IntuitionPosition+1