SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation
Abstract
While Text-to-Image (T2I) models have shown remarkablesuccess in generating photorealistic visual content, they still strugglewith the rigorous semantic alignment and logical reasoning requiredfor scientific imagery. Inspired by Peirce’s Semiotic Triad, we introduceScientific Image Reasoning (SciIR), a comprehensive resource for trainingand evaluation of scientific image generation. We formalize scientificreasoning into three core dimensions: Entity Structure (Icon), ScientificProcess (Index ), and Scientific Law (Symbol ). Specifically, to overcomethe scarcity of training data in scientific image generation, we elaboratelycreate SciIR-82k, a large-scale dataset containing over 80,000 high-qualityscientific image-text pairs from cutting-edge publications. The datasetis hierarchically organized according to the semiotic dimensions andincorporates a Scientific Reasoning Chain-of-Thought (Sci-RCoT) toexplicitly model underlying visual logic. For evaluation, we propose SciIR-Bench, which aligns with these three semiotic levels and employs anAtomic Checklist to convert the outcome-oriented scientific accuracyinto process-oriented, verifiable, fine-grained questions. Our extensiveexperiments reveal significant deficiencies in current models’ scientificreasoning capabilities. Furthermore, by fine-tuning on the SciIR-82kdataset, we developed the Qwen-Image-SciIR model, which achieves asubstantial improvement on the SciIR-Bench, increasing the final scorefrom 35% to 43%, laying a solid foundation for future advances in scientificimage generation.