Multiple Scale Latents for Learned Image Compression
Abstract
Most learned image compression systems rely on a single latent representation combined with a hyperprior, which limits their ability to efficiently capture image structure across spatial scales. In this work, we propose a hierarchical latent representation to improve the efficiency of the entropy model. Using multiple latents at different scales with their own entropy models, we aim to better capture the spatial structure of the latent representation.
Our experiments show that this approach achieves a -17.9% reduction in BD-rate over VVC on Kodak, demonstrating the effectiveness of multi-scale latent representations. Furthermore, the approach is orthogonal to other advances in learned image compression, making it a versatile addition to existing methods.
BibTeX
@article{brenig2026multiple,
title={Multiple Scale Latents for Learned Image Compression},
author={Jonas Brenig and Radu Timofte},
journal={Proceedings of the IEEE International Conference on Image Processing (ICIP)},
year={2026},
url={https://jbrenig.github.io/ICIP26-MSL}
}