Eight years in one setting
Nearly every idea in the field was first tried on labeled image classification, so its history is the history of the paradigms.
- 2018, posing the task. Dataset Distillation compressed MNIST’s 60,000 training images into 10 synthetic images that train a network to near-original accuracy with a few gradient steps, but only from a fixed initialization.
- 2020–2021, surrogates. Gradient matching and its augmentation-aware successor DSA trained networks from scratch on condensed sets. KIP solved the inner problem in closed form with neural tangent kernels. DM dropped the inner loop.
- 2022, the accuracy race. Trajectory matching matched long stretches of expert training. Parameterizations such as IDC and HaBa stored more images per byte, and IT-GAN and GLaD moved optimization into generator latent spaces.
- 2023, reaching ImageNet-1K. TESLA made trajectory matching fit in constant memory, reaching 50 images per class on ImageNet-1K on one GPU. SRe2L decoupled synthesis from the student entirely.
- 2024–2026, generative priors and a reckoning with labels. Diffusion models entered with Minimax Diffusion and D4M, and training-free guidance methods followed. At the same time A label is worth a thousand images and later re-evaluations showed how much of the reported accuracy comes from soft labels.
Reported milestones, and why they are not a ranking
Some of the numbers papers report, each taken from its own abstract:
| Paper | Reported result | Approach |
|---|---|---|
| SRe2L | 60.8% top-1, ImageNet-1K, 50 IPC | decoupled; relabels synthetic images with teacher soft labels |
| CDA | 63.2% top-1, ImageNet-1K, 50 IPC | decoupled; builds on SRe2L |
| RDED | 42% top-1 with ResNet-18, ImageNet-1K, 10 IPC | decoupled, realistic selection |
| EDC | 48.6% top-1 with ResNet-18, ImageNet-1K, 10 IPC | decoupled; design-space study |
| CoDA | 60.4% top-1, ImageNet-1K, 50 IPC | training-free text-to-image generation |
Evaluation recipes differ across these papers; for decoupled methods in particular, RD³ documents inconsistent post-evaluation protocols. The rows show how the achievable range moved, not which method is better. The evaluation page explains which variables to hold fixed before comparing.