Differentiable Hierarchical Visual Tokenization

Marius Aasan (University of Oslo) · Martine Hjelkrem Tan (University of Oslo) · Nico Catalano (Politecnico di Milano) · Changkyu Choi (UiT The Arctic University of Norway) · Adín Ramírez Rivera (University of Oslo)
backward-compatiblecompetitive performancedense-prediction tasksdifferentiable tokenizerend-to-end architecturefixed patch tokenshierarchical model selectionimage-level classificationinformation criteriapixel-level granularitypretrained modelsraster-to-vector conversionsemantic structurespatial structurevision transformers

Vision Transformers rely on fixed patch tokens that ignore the spatial and semantic structure of images. In this work, we introduce an end-to-end differentiable tokenizer that adapts to image content with pixel-level granularity while remaining backward-compatible with existing architectures for retrofitting pretrained models. Our method uses hierarchical model selection with information criteria to provide competitive performance in both image-level classification and dense-prediction tasks, and even supports out-of-the-box raster-to-vector conversion.