Generative Vision Atlas

Representation option

Pixels

Generate directly on raw image pixels; no learned tokenizer or latent space at all.

Used by (2)

JiT, Pixel MeanFlow

Alternatives on this axis

Continuous tokens, Deep-compression latent, Frozen encoder + trained decoder, Hybrid semantic + detail codebook, Multi-scale (next-scale) tokens, Semantic / foundation latent, VAE latent

← All Representation options