Condensing Large-Scale Datasets Directly with Minimal Information Loss
Abstract
Recent advancements in scaling dataset distillation rely heav-ily on decoupled information extraction pipelines, comprising Squeeze,Recover, and Relabel stages. Despite their scalability to large-scaledatasets, these methods suffer from prohibitive computational overheadand poor cross-architecture generalization. In this paper, we reveal theroot cause of these bottlenecks: the implicit dual-compression process,from data to model and back to images, inherently induces severe infor-mation loss. Crucially, we empirically and theoretically demonstrate thatthis loss creates a distribution shift that fundamentally compromises thewidely adopted Relabel strategy, transforming the pre-trained modelinto an unreliable labeler that yields sub-optimal labels. To overcomethese critical flaws, we propose CIM, a novel, metric-driven frameworkthat abandons the flawed dual-compression paradigm. Instead, CIM ex-plicitly quantifies and minimizes the information gap between the originaland synthetic datasets. By directly aligning the data distributions, ourapproach ensures high-fidelity information condensation and inherentlysatisfies the prerequisites for effective relabeling. Extensive experimentsdemonstrate that CIM establishes a new state-of-the-art. Notably, itdistills ImageNet-1K at an IPC =10 in merely 80 minutes on a singleRTX-4090 GPU, achieving an unprecedented 48.7% Top-1 accuracy onResNet-18 and significantly outperforming previous SOTA approaches,such as NRR-DD and DELT, by 2.6% and 2.9%, respectively. Our codeis available at https://github.com/LINs-lab/CIM.