arXiv · 2609.39657
Diffusable Latents from Structure-Agnostic Distillation
Abstract
Distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, enabling diffusion models to converge faster and reach higher sample quality. Standard distillation aligns the latent at each position to a co-located teacher feature, tying the latent layout to the teacher's. We show this constraint is unnecessary: aligning a single pooled image-level descriptor to the teacher's performs as well as or slightly better than dense position-wise distillation. We compare first-order and relational pooled objectives across latent shapes and teacher modalities. First-order matching extends naturally to 1D token-sequence latents and across modalities, where distilling a text encoder into an image autoencoder still improves diffusability; a relational objective based only on each image's nearest neighbours improves it as well. Code and blog post are available at https://github.com/AdrienRR/structure-agnostic-distillation and https://kyutai.org/blog/2026-09-28-structure-agnostic-distillation/.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Adrien Ramanana Rahary, Nicolas Dufour, Patrick Pérez, David Picard. 2026-09-30. Diffusable Latents from Structure-Agnostic Distillation. https://arxiv.org/abs/2609.39657
Cite the original work for its findings. Save a collection to share your selection of sources.