VisReason: A Large-Scale Dataset for Visual Chain-of-Thought Reasoning
Abstract
Chain-of-Thought (CoT) prompting has proven remarkablyeffective for eliciting complex reasoning in large language models (LLMs).Yet, its potential in multimodal large language models (MLLMs) re-mains largely untapped, hindered by the absence of large-scale datasetsthat capture the rich, spatially grounded reasoning intrinsic to visual un-derstanding. Existing visual-CoT resources are typically small, domain-specific, or lack the structured stepwise supervision necessary for com-positional visual reasoning. In this paper, we introduce VisReason, alarge-scale dataset designed to advance visual Chain-of-Thought rea-soning. VisReason comprises 489K annotated examples spanning fourdiverse domains, each featuring multi-round, RoI-grounded rationalesthat guide MLLMs through interpretable visual reasoning steps. Build-ing upon this, we curate VisReason-Pro, a 165K subset produced witha stronger GPT annotator, enriched with detailed reasoning traces anddepth-augmented spatial annotations derived from monocular depth andsegmentation cues. Fine-tuning strong MLLM backbones on VisRea-son and VisReason-Pro yields substantial improvements in step-by-stepvisual reasoning accuracy, RoI localization, interpretability, and fine-grained/spatial reasoning performance. These results demonstrate thatVisReason equips MLLMs with more systematic and verifiable visualreasoning capabilities. We envision VisReason as a cornerstone for cul-tivating human-like visual reasoning, paving the way toward the nextgeneration of multimodal intelligence.