Hyperbolic Hierarchical Clustering for Visual Representation Learning
Abstract
We investigate the token mixer in vision backbones by revisit-ing clustering, one of the most classic approaches in machine learning. Aneffective token mixer is a fundamental component of modern vision back-bones like vision Transformers, facilitating information exchange betweenimage patches. Mainstream token mixers, which rely on convolution, at-tention, MLP, or their hybrids, primarily focus on navigating the trade-offbetween accuracy and computational cost. However, a significant draw-back of these methods is their black-box nature; their encoding processis opaque and lacks interpretability. Diverging from these opaque designs,we introduce ClusterMixer, a transparent token mixer that is grounded ina clustering paradigm and interpretable by design. ClusterMixer explicitlyformulates the token mixing process through a hierarchical clusteringmechanism. To model the natural, tree-like relationships inherent in visualdata, the clustering is performed in hyperbolic space, which is well-suitedfor embedding hierarchies with low distortion. Building on this innovation,we present HCFormer, a new backbone architecture that integratesClusterMixer with a series of meticulously designed clustering strategiesto ensure robust performance across tasks. Extensive experiments demon-strate that HCFormer consistently outperforms its counterparts acrossdiverse tasks, including image classification, object detection, instancesegmentation, and semantic segmentation. Considering its transparencyand efficacy, we hope HCFormer can facilitate a paradigm shift towardinterpretable backbones.