paper-with-me

홈 › Papers

ScImage: How Good Are Multimodal Large Language Models at Scientific Text-to-Image Generation?

2024-12-03 · Leixin Zhang, Steffen Eger, Yinjie Cheng, Weihe Zhai, Jonas Belouadi, Christoph Leiter, Simone Paolo Ponzetto, Fahimeh Moafian, Zhixue Zhao

Multimodal large language models (LLMs) have demonstrated impressive capabilities in generating high-quality images from textual instructions. However, their performance in generating scientific images--a critical application for accelerating scientific progress--remains underexplored. In this work, we address this gap by introducing ScImage, a benchmark designed to evaluate the multimodal capabilities of LLMs in generating scientific images from textual descriptions. ScImage assesses three key dimensions of understanding: spatial, numeric, and attribute comprehension, as well as their combinations, focusing on the relationships between scientific objects (e.g., squares, circles). We evaluate five models, GPT-4o, Llama, AutomaTikZ, Dall-E, and StableDiffusion, using two modes of output generation: code-based outputs (Python, TikZ) and direct raster image generation. Additionally, we examine four different input languages: English, German, Farsi, and Chinese. Our evaluation, conducted with 11 scientists across three criteria (correctness, relevance, and scientific accuracy), reveals that while GPT-4o produces outputs of decent quality for simpler prompts involving individual dimensions such as spatial, numeric, or attribute understanding in isolation, all models face challenges in this task, especially for more complex prompts.

📄 PDF Abstract BibTeX arXiv:2412.02368

Code (1)

leixin-zhang/scimage 공식 구현

Tasks

AttributeImage GenerationText to Image GenerationText-to-Image Generation

Similar Papers 제목 키워드 기반

Transforming Science with Large Language Models: A Survey on AI-assisted Scientific Discovery, Experimentation, Content Generation, and Evaluation

2025-02-07 · Steffen Eger, Yong Cao, Jennifer D'Souza, Andreas Geiger 외

With the advent of large multimodal language models, science is now at a threshold of an AI-based technological transformation. Recently, a plethora of new AI models and tools has been proposed, promising to empower rese…

scientific discoverySurvey

SCITUNE: Aligning Large Language Models with Scientific Multimodal Instructions

2023-07-03 · Sameera Horawalavithana, Sai Munikoti, Ian Stewart, Henry Kvinge

Instruction finetuning is a popular paradigm to align large language models (LLM) with human intent. Despite its popularity, this idea is less explored in improving the LLMs to align existing foundation models with scien…

S1-MMAlign: A Large-Scale, Multi-Disciplinary Dataset for Scientific Figure-Text Understanding

2026-01-01 · He Wang, Longteng Guo, Pengkang Huo, Xuanxu Lin 외 arxiv

Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. We pre…

A unified multimodal understanding and generation model for cross-disciplinary scientific research

2026-01-04 · Xiaomeng Yang, Zhiyu Tan, Xiaohui Zhong, Mengping Yang 외 arxiv

Scientific discovery increasingly relies on integrating heterogeneous, high-dimensional data across disciplines nowadays. While AI models have achieved notable success across various scientific domains, they typically re…

Visual Question AnsweringWeather Forecasting

Innovator-VL: A Multimodal Large Language Model for Scientific Discovery

2026-01-27 · Zichen Wen, Boxue Yang, Shuang Chen, Yaojie Zhang 외 arxiv

We present Innovator-VL, a scientific multimodal large language model designed to advance understanding and reasoning across diverse scientific domains while maintaining excellent performance on general vision tasks. Con…

Reinforcement LearningMultimodal Reasoning