Generative Vision Atlas

Representation option

VAE latent

A compact continuous latent from a KL-regularized autoencoder (e.g. SD-VAE, 4-16 channels), trained purely for pixel reconstruction.

Introduced by

Used by (5)

DiT, LDM / Stable Diffusion, REPA-E, SD3 / MMDiT, VA-VAE / LightningDiT

Alternatives on this axis

Continuous tokens, Deep-compression latent, Frozen encoder + trained decoder, Hybrid semantic + detail codebook, Multi-scale (next-scale) tokens, Pixels, Semantic / foundation latent

← All Representation options