When Does Closeness in Distribution Imply Representational Similarity? An Identifiability Perspective

Emanuele Marconato (University of Trento, Via Calepina 14, 38122, Trento VAT: IT00340520220) · Beatrix Nielsen (Technical University of Denmark, IT University of Copenhagen) · Andrea Dittadi (Helmholtz AI & TU Munich) · Luigi Gresele (University of Copenhagen)
autoregressive language modelscifar-10data likelihooddissimilar representationsdistributional distanceempirical validationidentifiability theorykullback--leibler divergencemodel distributionsmodel familypre-training approachesrepresentational invariancerepresentational similaritysimilarity metricssynthetic experimentswider networks

When and why representations learned by different deep neural networks are similar is an active research topic. We choose to address these questions from the perspective of identifiability theory, which suggests that a measure of representational similarity should be invariant to transformations that leave the model distribution unchanged. Focusing on a model family which includes several popular pre-training approaches, e.g., autoregressive language models, we explore when models which generate distributions that are close have similar representations. We prove that a small Kullback--Leibler divergence between the model distributions does not guarantee that the corresponding representations are similar. This has the important corollary that models with near-maximum data likelihood can still learn dissimilar representations