When the Teacher Has More Bits: Self-Teacher Latent Distillation for Learned Image Compression
Abstract
Learned image compression (LIC) operates under a rate–distortion (RD) trade-off, where representation quality is constrained bythe target bitrate. We revisit knowledge distillation in LIC through thelens of bitrate asymmetry and introduce a self-teacher distillation frame-work, where a high-rate instance of a codec supervises multiple lower-rateencoders of identical architecture. Because allocating more bits naturallyleads to richer latent representations, the high-rate model provides infor-mative supervision across rate levels. Direct latent matching, however,is problematic under tight rate budgets. We therefore propose variance-normalized latent distillation (VNLD), a rate-aware alignment strategythat scales channel-wise supervision by the teacher’s variance, selectivelytransferring stable, informative structure while suppressing componentsthat cannot be reliably reproduced at lower rates. Across different dis-tillation objectives, networks, and bitrate levels, self-teacher distillation,particularly with VNLD, improves RD performance and yields consistentBD-rate gains over RD-only training. Our method remains compatiblewith fixed-decoder deployments, such as those targeted by JPEG AIstandards.