Living research atlas · updated every 3 days
Follow the representation.
Modern image generation has mostly settled its arguments about objectives and architectures. What it has not settled is where generation should happen: in a compression code, inside a foundation model's features, in discrete tokens, or in raw pixels.
This atlas is organized around that question and the others like it. Every paper here sits inside a research line defined by one explicit bet, so you can see what a group is claiming, what it costs them, and who disagrees.
New to this? Start with the fundamentals
Stanford CME296: Diffusion & Large Vision Models ↗
This atlas begins where a course ends. It assumes you already know how diffusion and flow matching work, and spends its time on what the field is still arguing about. If you want that foundation first, or want it taught properly rather than reconstructed from papers, this lecture course is the recommended starting point.
Free on YouTube · then come back and read the representation debate with the background to judge it.
The argument the field is having
Five live answers to one question, ordered by how much meaning the latent carries. Each is an active research line with its own evidence and its own weaknesses.
The honest summary: at their best-tuned settings all five land within roughly the same range on ImageNet, and the published numbers cannot rank them, because guidance method, model size and training budget all move at once. See why →
How this is organized
Six design questions. Every research line answers exactly one of them, and lines answering the same question compete directly.
What space do you generate in?
What do you train the model to predict?
What shapes it beyond the objective?
How does the prompt reach the generator?
What happens at sampling time?
How are understanding and generation arranged?
What this does that a paper list does not
Refuses to publish a leaderboard
One paper reports the same model at 1.65, 1.49, 1.14 and 1.06 under four guidance methods. That swing is bigger than most architectural claims. Every number here carries the guidance, budget and model size that produced it.
Tracks what is unsolved
11 open problems, each linked to the papers attacking it and the research lines that carry it as a known weakness. Where a line is weak, its page says so.
Updates itself
A pipeline sweeps arXiv and Hugging Face every three days, filters for genuine relevance, and proposes candidates for review. Nothing is added automatically and nothing is claimed without a fetched source.
Separate domain
Medical imaging
Kept apart from the rest of the atlas on purpose. Clinical imaging answers the same design question differently, and it answers to a different evidence bar: not a better distribution distance, but whether a diagnostic model improves and whether a radiologist is fooled.
27 papers across 5 lines, organized by which space they generate in and why that space was chosen →
Start here
You want the fundamentals first ↗
Stanford CME296: Diffusion & Large Vision Models. The background this atlas assumes, taught in order rather than reconstructed from papers.
New to the latent debate
The line that asks whether the generative space should be a pretrained representation at all.
You want the numbers
Head to head on ImageNet and text-to-image, with what the numbers can and cannot support.
You want the story in order
How the field moved, and how long ideas sat unused before something made them work.
You want to run something
Four notebooks with real outputs from an A100: what each latent discards, RAE against the VAE it replaces, how editing methods differ, and why guidance breaks benchmarks.
You want to browse by subject
Six topic areas plus medical, each with its papers, its research lines, and its open problems.