paper-with-me

Papers

SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation

2026-06-29 · Zhiyuan Ma, Zhengfeng Shi, Yuning An, Peize Li, Jiabao Wei, Ruijie Li, Junhao Xiao, Jianjun Li, Bowen Zhou hf

While Text-to-Image (T2I) models have shown remarkable success in generating photorealistic visual content, they still struggle with the rigorous semantic alignment and logical reasoning required for scientific imagery. Inspired by Peirce's Semiotic Triad, we introduce Scientific Image Reasoning (SciIR), a comprehensive resource for training and evaluation of scientific image generation. We formalize scientific reasoning into three core dimensions: Entity Structure (Icon), Scientific Process (Index), and Scientific Law (Symbol). Specifically, to overcome the scarcity of training data in scientific image generation, we elaborately create SciIR-82k, a large-scale dataset containing over 80,000 high-quality scientific image-text pairs from cutting-edge publications. The dataset is hierarchically organized according to the semiotic dimensions and incorporates a Scientific Reasoning Chain-of-Thought (Sci-RCoT) to explicitly model underlying visual logic. For evaluation, we propose SciIR-Bench, which aligns with these three semiotic levels and employs an Atomic Checklist to convert the outcome-oriented scientific accuracy into process-oriented, verifiable, fine-grained questions. Our extensive experiments reveal significant deficiencies in current models' scientific reasoning capabilities. Furthermore, by fine-tuning on the SciIR-82k dataset, we developed the Qwen-Image-SciIR model, which achieves a substantial improvement on the SciIR-Bench, increasing the final score from 35\% to 43\%, laying a solid foundation for future advances in scientific image generation.

📄 PDF Abstract BibTeX arXiv:2606.30124

Code (0)

등록된 구현이 없습니다.

Tasks

Logical ReasoningImage Generation

Similar Papers 제목 키워드 기반

CNVid-3.5M: Build, Filter, and Pre-Train the Large-Scale Public Chinese Video-Text Dataset

2023-01-01 · CVPR 2023 1 · Tian Gan, Qing Wang, Xingning Dong, Xiangyuan Ren 외

Owing to well-designed large-scale video-text datasets, recent years have witnessed tremendous progress in video-text pre-training. However, existing large-scale video-text datasets are mostly English-only. Though th…

Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark

2022-02-14 · Jiaxi Gu, Xiaojun Meng, Guansong Lu, Lu Hou 외

Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale dat…

BenchmarkingContrastive Learningimage-classificationImage Classification+6

BigDetection: A Large-scale Benchmark for Improved Object Detector Pre-training

2022-03-24 · Likun Cai, Zhi Zhang, Yi Zhu, Li Zhang 외

Multiple datasets and open challenges for object detection have been introduced in recent years. To build more general and powerful object detection systems, in this paper, we construct a new large-scale benchmark termed…

Objectobject-detectionObject Detection

Large-scale Unsupervised Semantic Segmentation

2021-06-06 · ShangHua Gao, Zhong-Yu Li, Ming-Hsuan Yang, Ming-Ming Cheng 외

Empowered by large datasets, e.g., ImageNet, unsupervised learning on large-scale data has enabled significant advances for classification tasks. However, whether the large-scale unsupervised semantic segmentation can be…

DiversityRepresentation LearningSegmentationSemantic Segmentation+1

CCMB: A Large-scale Chinese Cross-modal Benchmark

2022-05-08 · Chunyu Xie, Heng Cai, Jincheng Li, Fanjing Kong 외

Vision-language pre-training (VLP) on large-scale datasets has shown premier performance on various downstream tasks. In contrast to plenty of available benchmarks with English corpus, large-scale pre-training datasets a…

image-classificationImage ClassificationImage GenerationImage Retrieval+9