Survey
Sachdeva & McAuley survey
Data Distillation: A Survey
Noveen Sachdeva, Julian McAuley
TMLR 2023 · first public 2023-01-11 · arXiv 2301.04272
In one paragraph
This survey presents a formal framework for data distillation with a detailed taxonomy of existing approaches, and covers the method across three data modalities: images, graphs, and user-item interactions (recommender systems), identifying current challenges and future research directions for each.
Where it sits
- Setting: Image classification
- Setting: Graphs
- Setting: Other data types
Abstract (verbatim from arXiv)
The popularity of deep learning has led to the curation of a vast number of massive and multifarious datasets. Despite having close-to-human performance on individual tasks, training parameter-hungry models on large datasets poses multi-faceted problems such as (a) high model-training time; (b) slow research iteration; and (c) poor eco-sustainability. As an alternative, data distillation approaches aim to synthesize terse data summaries, which can serve as effective drop-in replacements of the original dataset for scenarios like model training, inference, architecture search, etc. In this survey, we present a formal framework for data distillation, along with providing a detailed taxonomy of existing approaches. Additionally, we cover data distillation approaches for different data modalities, namely images, graphs, and user-item interactions (recommender systems), while also identifying current challenges and future research directions.
BibTeX (generated; prefer the venue's official entry)
@article{sachdeva2023data,
title = {Data Distillation: A Survey},
author = {Noveen Sachdeva and Julian McAuley},
journal = {TMLR 2023},
year = {2023}
}