Research line · emerging
Medical: foundation-model latent
Generate inside a medical foundation model's feature space, the RAE idea applied to clinical images.
What defines membership
A medical vision foundation model's representation is a better generative substrate than either a borrowed natural-image latent or a reconstruction-trained domain latent.
How the line developed
Read top to bottom: what came before, the idea itself, the evidence for it, what improved, and where it breaks.
The idea
STREAM · 2026-06core
Riemannian flow matching directly inside a histopathology foundation model's patch-token space, with an anisotropic decoder; motivated by the conditioning collapse that appears when foundation features only supply a condition.
Limitation
Retinal FM tokenizers · 2026-08core
Four retinal foundation models tested as generative latents: the advantage over conventional latent diffusion largely vanishes when judged by classifiers trained on real images.
What it gets right
- Inherits the semantics of a model already trained on the target anatomy
- Directly addresses the unvalidated assumption underneath the borrowed-VAE line
- Medical foundation models already exist and are widely used for understanding
- The two papers come from unrelated groups, one UK academic and one Korean industry, arriving at the same idea independently within months of each other
Where it is weak
- Two papers old, both from 2026. Searched deliberately in September 2026 across arXiv and PubMed for a third group generating inside a medical foundation model's own representation space, and found none. This is recorded as a genuinely nascent area rather than as incomplete coverage
- The retinal result suggests gains may be partly an artifact of evaluating with the same foundation model that defines the latent
- Untouched outside pathology and retina: no chest X-ray or general radiology work
Competing answers
Open problems it has not solved
- Most medical image generation runs inside an autoencoder trained on natural photographs, and nobody has tested whether that latent preserves clinically relevant detail.
- Semantic / foundation-model latents discard much of the high-frequency pixel detail (exact color, texture, fine structure) that faithful reconstruction — and, later, edit-region preservation — depends on.