Representation option
Frozen encoder + trained decoder
Keep a pretrained vision encoder entirely frozen as the latent space; train only a lightweight decoder to map that space back to pixels.
Introduced by
- RAE 2025-10
Alternatives on this axis
Continuous tokens, Deep-compression latent, Hybrid semantic + detail codebook, Multi-scale (next-scale) tokens, Pixels, Semantic / foundation latent, VAE latent