GRADE: Benchmarking Discipline-Informed Reasoning in Image Editing
Abstract
Unified multimodal models target joint understanding, rea-soning, and generation, but current image editing benchmarks are largelyconfined to natural images and shallow commonsense reasoning, offeringlimited assessment of this capability under structured, domain-specificconstraints. In this work, we introduce GRADE, the first benchmarkto assess discipline-informed knowledge and reasoning in image editing.GRADE comprises 520 carefully curated samples across 10 academicdomains, spanning from natural science to social science. To supportrigorous evaluation, we propose a multi-dimensional evaluation proto-col that jointly assesses Discipline Reasoning, Visual Consistency, andLogical Readability. Extensive experiments on 20 state-of-the-art open-source and closed-source models reveal substantial limitations in currentmodels under implicit, knowledge-intensive editing settings, leading tolarge performance gaps. Beyond quantitative scores, we conduct rigor-ous analyses and ablations to expose model shortcomings and identifythe constraints within disciplinary editing. Together, GRADE pinpointskey directions for the future development of unified multimodal models,advancing the research on discipline-informed image editing and reason-ing. Our benchmark and evaluation code are publicly released.