Generative Vision Atlas

Representation option

Semantic / foundation latent

A latent taken from (or aligned to) a pretrained vision foundation encoder's feature space, so it is high-dimensional and semantically structured rather than reconstruction-only.

Introduced by

Used by (10)

BLIP3-o, Emu2, MAETok, MetaQuery, RAE, RAEv2, Scale-RAE, SVG, SVG-T2I, unCLIP / DALL-E 2

Alternatives on this axis

Continuous tokens, Deep-compression latent, Frozen encoder + trained decoder, Hybrid semantic + detail codebook, Multi-scale (next-scale) tokens, Pixels, VAE latent

← All Representation options