Architecture option
Diffusion Transformer (DiT)
A plain Vision-Transformer backbone for diffusion/flow models, operating on a grid of latent patch tokens with adaptive layer norm for timestep/class conditioning.
Introduced by
- DiT 2022-12
Alternatives on this axis
DiT with diffusion head (DiT^DH), MM-DiT (dual-stream joint attention), Single-stream (unified-weight) DiT, UNet backbone