Generative Vision Atlas

Timeline

Every paper in the general-domain atlas by date, with the line it belongs to. Medical imaging has its own separate timeline of ideas. The point of reading it in order is to see how long an idea sits unused before something makes it work: generating in a semantic space was tried in 2022 and abandoned, then became the most active question in the field three years later.

2021

Before the shift. Latent diffusion and masked autoencoders arrive separately; nobody yet connects them.

02
CLIPlandmark · Contrastive language-image pretraining
05
CDMlandmark · Cascaded and multiscale pixel diffusion
11
MAElandmark · Self-distillation representations
12
LDM / Stable Diffusionlandmark · VAE-latent diffusion

2022

Latent diffusion becomes the default, and CLIP-space generation is tried once and set aside.

04
Flamingolandmark · Cross-attention into a frozen language model
04
unCLIP / DALL-E 2landmark · Representation-space generation
05
Imagenlandmark
07
Classifier-Free Guidancelandmark · Guidance and sampling
08
Prompt-to-Promptlandmark · Training-free attention manipulation
09
Re-Imagenlandmark · Agentic and search-augmented generation
09
Rectified Flowlandmark · Flow matching and rectified flow
10
Flow Matchinglandmark · Flow matching and rectified flow
11
InstructPix2Pixlandmark · Editing in the VAE latent
11
Plug-and-Play · Training-free attention manipulation
12
DiTlandmark · VAE-latent diffusion

2023

The backbone becomes a transformer, and self-supervised features get good enough to matter.

01
BLIP-2landmark · Cross-attention into a frozen language model
01
I-JEPA · Self-distillation representations
01
simple diffusion · Cascaded and multiscale pixel diffusion
02
ControlNetlandmark · Adapter-based conditioning
02
T2I-Adapter · Adapter-based conditioning
03
SigLIP · Contrastive language-image pretraining
04
DINOv2landmark · Self-distillation representations
04
LLaVAlandmark · Encoder plus projector
04
MasaCtrl · Training-free attention manipulation
05
Otter · Cross-attention into a frozen language model
06
MagicBrush · Editing in the VAE latent
08
IP-Adapterlandmark · Adapter-based conditioning
09
ViT Registers · Self-distillation representations
10
Idea2Img · Agentic and search-augmented generation
11
Emu Edit · Editing in the VAE latent
12
Emu2 · Representation-space generation, Unified understanding and generation
12
GIVT · Continuous-token autoregression
12
AM-RADIO · Agglomerative multi-teacher distillation

2024

Flow matching wins the objective argument. Autoregression mounts a serious challenge. Alignment losses appear.

01
HDiT · Cascaded and multiscale pixel diffusion
01
InstantID · Adapter-based conditioning
01
SiT · Flow matching and rectified flow, VAE-latent diffusion
03
SD3 / MMDiTlandmark · Flow matching and rectified flow, VAE-latent diffusion
04
Guidance interval · Guidance and sampling
04
PuLID · Adapter-based conditioning
04
VARlandmark · Discrete-token autoregression
05
Chameleonlandmark · Unified understanding and generation
06
Autoguidancelandmark · Guidance and sampling
06
CFG++ · Guidance and sampling
06
MARlandmark · Continuous-token autoregression
07
GenArtist · Agentic and search-augmented generation
07
Theia · Agglomerative multi-teacher distillation
08
LLaVA-OneVision · Encoder plus projector
08
Transfusionlandmark · Continuous-token autoregression, Unified understanding and generation
08
UNIC · Agglomerative multi-teacher distillation
09
Emu3 · Unified understanding and generation
09
NVLM · Cross-attention into a frozen language model
09
Qwen2-VL · Native-resolution vision-language models
10
Adaptive Projected Guidance · Guidance and sampling
10
DC-AE · VAE-latent diffusion
10
Fluidlandmark · Continuous-token autoregression
10
Janus · Unified understanding and generation
10
REPAlandmark · Representation-aligned latents
10
RF-Inversion · Inversion for flow models
10
Shortcut Models · Natively few-step objectives
10
SiD2landmark · Cascaded and multiscale pixel diffusion
11
AIMv2 · Contrastive language-image pretraining
11
Edify Image · Cascaded and multiscale pixel diffusion
11
OminiControllandmark · Adapter-based conditioning, In-context editing
11
RF-Solver · Inversion for flow models
11
Stable Flow · Training-free attention manipulation
12
FireFlow · Inversion for flow models
12
RADIOv2.5 · Agglomerative multi-teacher distillation
12
TarFlowlandmark · Normalizing-flow revival
12
TokenFlow · Semantic-plus-detail hybrids

2025

The latent itself becomes the contested question: align it, replace it, or delete it.

01
ACE++ · In-context editing
01
Janus-Pro · Unified understanding and generation
01
VA-VAE / LightningDiT · Representation-aligned latents
02
KV-Edit · Training-free attention manipulation
02
MAETok · Representation-aligned latents
02
Qwen2.5-VL · Native-resolution vision-language models
02
SigLIP 2 · Contrastive language-image pretraining
03
ICE-Bench · Evaluation and benchmarks
03
IMM · Natively few-step objectives
03
Lumina-Image 2.0 · VAE-latent diffusion
03
SANA-Sprint · Natively few-step objectives
04
GigaTok · Discrete-token autoregression
04
ICEdit · In-context editing
04
InternVL3 · Native-resolution vision-language models
04
MetaQuery · Encoder plus projector, Unified understanding and generation
04
Perception Encoder · Contrastive language-image pretraining
04
PixelFlow · Cascaded and multiscale pixel diffusion, Flow matching and rectified flow
04
REPA-E · Representation-aligned latents
04
RISEBench · Evaluation and benchmarks
04
Step1X-Edit · Editing in the VAE latent, In-context editing
04
Web-SSL · Self-distillation representations
05
BAGEL · Editing inside a unified model, Unified understanding and generation
05
BLIP3-o · Unified understanding and generation
05
Flow-GRPO · Reinforcement learning and preference alignment
05
HiDream-I1 · VAE-latent diffusion
05
ImgEdit · Evaluation and benchmarks
05
KRIS-Bench · Evaluation and benchmarks
05
MeanFlowlandmark · Natively few-step objectives
05
UniEval · Evaluation and benchmarks
06
FLUX.1 Kontextlandmark · Editing in the VAE latent, In-context editing
06
OmniGen2 · Editing inside a unified model, In-context editing, Unified understanding and generation
06
RefEdit
06
Show-o2 · Unified understanding and generation
06
STARFlow · Normalizing-flow revival
06
Transition Matching · Transition matching
06
V-JEPA 2 · Self-distillation representations
07
PixNerdlandmark · Single-stage pixel transformers
07
REG · Representation-aligned latents
07
Resurrect MAR · Continuous-token autoregression
08
DINOv3 · Self-distillation representations
08
NextStep-1 · Continuous-token autoregression
08
Qwen-Image · VAE-latent diffusion
09
DiffusionNFT · Reinforcement learning and preference alignment
09
EditScore · Reinforcement learning and preference alignment
09
GLASS Flows · Guidance and sampling
09
HunyuanImage 3.0 · Unified understanding and generation
09
Seedream 4.0 · In-context editing, VAE-latent diffusion
09
Tokenizer Post-Training · Discrete-token autoregression
10
Emu3.5 · Editing inside a unified model, Unified understanding and generation
10
Neon · Reinforcement learning and preference alignment
10
There is No VAE · Single-stage pixel transformers
10
Pico-Banana-400K
10
RAElandmark · Representation-space generation
10
Rectified-CFG++ · Guidance and sampling
10
SVG · Representation-space generation
10
Demystifying TM · Transition matching
10
VFM-VAE · Representation-space generation
11
DiP · Single-stage pixel transformers
11
JiT · Single-stage pixel transformers
11
PixelDiT · Single-stage pixel transformers
11
Qwen3-VL · Native-resolution vision-language models
12
GenEval 2 · Evaluation and benchmarks
12
PS-VAE · Editing in a representation latent
12
REPA spatial-structure study · Representation-aligned latents
12
SVG-T2I · Representation-space generation
12
TM Design Space · Transition matching

2026

Consolidation and new paradigms. Hybrids soften the VAE-versus-representation split while one-step objectives attack from a different direction.

01
C-RADIOv4 · Agglomerative multi-teacher distillation
01
DiDAE · Editing in a representation latent
01
Pixel MeanFlow · Natively few-step objectives, Single-stage pixel transformers
01
Scale-RAE · Representation-space generation
01
TM Distillation · Transition matching
02
Drifting Modelslandmark · Natively few-step objectives
02
FlatDINO · Semantic-plus-detail hybrids
02
Latent Forcing · Semantic-plus-detail hybrids
02
LV-RAE · Semantic-plus-detail hybrids
02
Masked Bit Modeling · Discrete-token autoregression
03
PixelREPA · Single-stage pixel transformers, Representation-space generation
03
RPiAE · Editing in a representation latent
05
DecQ · Semantic-plus-detail hybrids
05
HiDream-O1-Image · Editing inside a unified model, Single-stage pixel transformers
05
HyperDiT · Single-stage pixel transformers
05
PAE · Semantic-plus-detail hybrids
05
Qwen-Image 2.0 · VAE-latent diffusion
05
RAEv2 · Representation-space generation
05
Register Guidance · Single-stage pixel transformers
05
STARFlow2 · Normalizing-flow revival, Unified understanding and generation
06
CrossFlow · Semantic-plus-detail hybrids, Natively few-step objectives
06
Distilling Drifting Transformers · Natively few-step objectives, Representation-space generation
06
Latent Diffusability · Semantic-plus-detail hybrids
06
Parallel Rollout Approximation · Single-stage pixel transformers
07
Chimera · VAE-latent diffusion
07
dRAE · Representation-space generation, Discrete-token autoregression
07
Parallel Decoding Distillation · Natively few-step objectives
07
Pixel-space survey · Single-stage pixel transformers
07
SearchGen · Agentic and search-augmented generation
07
Self-Sample Guidance · Single-stage pixel transformers, Guidance and sampling
07
Three-Body Scattering
08
Any-OPD · Natively few-step objectives
08
ContextBias · Evaluation and benchmarks
08
GenFirst · Semantic-plus-detail hybrids
08
MOSAIK · Single-stage pixel transformers
08
Observation Operators · Single-stage pixel transformers
08
Pixel T2I empirical study · Single-stage pixel transformers
08
Second-Order Drifting · Natively few-step objectives
08
UniSpace · Editing in a representation latent
09
Agentic Visual Generation · Agentic and search-augmented generation
09
LLaDA-Image · VAE-latent diffusion
09
PixSGR · Single-stage pixel transformers