Which model does what
The question most people arrive with is whether a given model can both generate and edit, and what it is working on top of. This table answers both. The basis column is the space the model actually operates in, which is what determines its failure modes.
10 of 40 systems here do both generation and editing. 6 also do image understanding in the same model.
Almost all of them sit on a VAE latent. The exception worth knowing is HiDream-O1-Image, which drops the autoencoder entirely and embeds raw pixel patches, text and condition tokens into one shared transformer, so editing and generation are literally the same forward pass.
| Model | Generate | Edit | Multi-ref | Understand | Tokenize | Built on | Availability | Date |
|---|---|---|---|---|---|---|---|---|
| LLaDA-Image | ✓ | ✓ | · | ✓ | · | VAE latent | open weights | 2026-09 |
| UniSpace | ✓ | ✓ | · | ✓ | · | not recorded | unclear | 2026-08 |
| HiDream-O1-Image | ✓ | ✓ | · | ✓ | · | Pixel space | open weights | 2026-05 |
| Qwen-Image 2.0 | ✓ | ✓ | · | · | · | VAE latent | unclear | 2026-05 |
| RAEv2 | · | · | · | · | ✓ | Representation latent | code only | 2026-05 |
| Drifting Models | ✓ | · | · | · | · | not recorded | code only | 2026-02 |
| Scale-RAE | ✓ | · | · | · | ✓ | Representation latent | open weights | 2026-01 |
| SVG-T2I | ✓ | · | · | · | ✓ | Representation latent | open weights | 2025-12 |
| JiT | ✓ | · | · | · | · | Pixel space | open weights | 2025-11 |
| Emu3.5 | ✓ | ✓ | · | ✓ | · | not recorded | open weights | 2025-10 |
| RAE | · | · | · | · | ✓ | Representation latent | open weights | 2025-10 |
| SVG | · | · | · | · | ✓ | Representation latent | open weights | 2025-10 |
| HunyuanImage 3.0 | ✓ | · | · | ✓ | · | not recorded | open weights | 2025-09 |
| Seedream 4.0 | ✓ | ✓ | ✓ | · | · | VAE latent | API only | 2025-09 |
| NextStep-1 | ✓ | · | · | · | · | Continuous tokens | open weights | 2025-08 |
| Qwen-Image | ✓ | ✓ | · | · | · | VAE latent | open weights | 2025-08 |
| FLUX.1 Kontext | ✓ | ✓ | ✓ | · | · | not recorded | open weights | 2025-06 |
| OmniGen2 | ✓ | ✓ | · | ✓ | · | not recorded | open weights | 2025-06 |
| Show-o2 | ✓ | · | · | ✓ | · | not recorded | open weights | 2025-06 |
| STARFlow | ✓ | · | · | · | · | not recorded | unclear | 2025-06 |
| BAGEL | ✓ | ✓ | · | ✓ | · | not recorded | open weights | 2025-05 |
| BLIP3-o | ✓ | · | · | ✓ | · | not recorded | open weights | 2025-05 |
| HiDream-I1 | ✓ | · | · | · | · | VAE latent | open weights | 2025-05 |
| MeanFlow | ✓ | · | · | · | · | not recorded | code only | 2025-05 |
| ICEdit | · | ✓ | · | · | · | not recorded | open weights | 2025-04 |
| MetaQuery | ✓ | · | · | ✓ | · | not recorded | code only | 2025-04 |
| Step1X-Edit | · | ✓ | · | · | · | not recorded | open weights | 2025-04 |
| Lumina-Image 2.0 | ✓ | · | · | · | · | VAE latent | open weights | 2025-03 |
| ACE++ | · | ✓ | · | · | · | not recorded | open weights | 2025-01 |
| Janus-Pro | ✓ | · | · | ✓ | · | not recorded | open weights | 2025-01 |
| MAR | ✓ | · | · | · | · | Continuous tokens | open weights | 2024-06 |
| VAR | ✓ | · | · | · | · | Discrete tokens | open weights | 2024-04 |
| SD3 / MMDiT | ✓ | · | · | · | · | VAE latent | open weights | 2024-03 |
| SiT | ✓ | · | · | · | · | VAE latent | open weights | 2024-01 |
| Emu2 | ✓ | · | · | ✓ | · | Representation latent | open weights | 2023-12 |
| Emu Edit | · | ✓ | · | · | · | not recorded | closed | 2023-11 |
| DiT | ✓ | · | · | · | · | VAE latent | open weights | 2022-12 |
| InstructPix2Pix | · | ✓ | · | · | · | not recorded | open weights | 2022-11 |
| unCLIP / DALL-E 2 | ✓ | · | · | · | · | Representation latent | closed | 2022-04 |
| LDM / Stable Diffusion | ✓ | · | · | · | · | VAE latent | open weights | 2021-12 |
A blank basis means the paper is a method or analysis contribution rather than a system built on one particular latent. Medical models are listed separately under medical imaging.