Generative Vision Atlas

Research line · emerging

Normalizing-flow revival

Return to exact-likelihood invertible models, scaled up with transformers and trained in a latent space.

What defines membership

Exact likelihoods and invertibility are worth having, and the old scaling barriers were architectural rather than fundamental.

How the line developed

Read top to bottom: what came before, the idea itself, the evidence for it, what improved, and where it breaks.

The idea

TarFlow · 2024-12landmark

Transformer autoregressive flow blocks with alternating direction, noise augmentation and flow-specific guidance. The architecture the line is built on.

Improvement

STARFlow2 · 2026-05strong-followup

Fuses the flow stream with a vision-language stream so visual tokens enter the same key-value cache as text.

At scale

STARFlow · 2025-06core

Scales TarFlow to high resolution in a pretrained autoencoder's latent, with an explicit ablation preferring latent over pixel modelling.

What it gets right

  • Exact likelihoods, which diffusion and flow matching cannot provide
  • Invertible by construction, which matters for editing and inverse problems
  • Backed by a major industrial lab rather than a single academic result

Where it is weak

  • Far less validated than diffusion at frontier scale
  • Architecturally constrained by invertibility
  • Thin follow-up literature so far
  • Every paper in this line is from one group at Apple. Unlike the other lines here, the atlas found no independent replication, which is itself the clearest evidence of how narrow the revival currently is

Competing answers